Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plaus
Record details
Published: 4 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Healthcare · Agents & autonomy
Retrieved: 5 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
arXiv cs.CY · 5 August 2026
Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation
arXiv cs.HC · 6 August 2026
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations
arXiv cs.CY · 7 August 2026
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
arXiv cs.CY · 30 July 2026
Hearsay: Vision-Language Medical Diagnoses Without an Image
arXiv cs.CY · 30 July 2026
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study
JMIR (Journal of Medical Internet Research) · 29 July 2026
How to cite this record
ethics.ai (4 August 2026), “Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems,” evidence record 16553, https://ethics.ai/record/16553 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.