Evidence record 1233 · automatically gathered

The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge

Multi-agent debate systems are typically evaluated only on whether the final answer is correct, overlooking the quality of the intermediate reasoning that debate is designed to produce. This paper studies the relationship between three signals in multi-agent debate: token-level log-probability distributions over reasoning tokens, LLM-as-judge rubric scores assigned to those tokens, and final task accuracy. We examine whether internal confidence signals predict externally evaluated reasoning qual

Record details

Published: 9 June 2026
Source: arXiv
Category: Research
Topics: Healthcare · Agents & autonomy
Retrieved: 14 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (9 June 2026), “The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge,” evidence record 1233, https://ethics.ai/record/1233 (originally published by arXiv).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.