Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced sycophancy. We test this attribution across four model families and find it largely wrong: pretrained base models exhibit the same substitution pattern as their Instruct variants, averaging higher yield than Instruct. Using activation patching, we localize the corruption to a narrow mid-layer window where attention carr
Record details
Published: 13 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Position: Assistive Agents Need Accessibility Alignment
arXiv · 13 May 2026
Unweighted ranking for value-based decision making with uncertainty
arXiv · 13 May 2026
No More, No Less: Task Alignment in Terminal Agents
arXiv · 12 May 2026
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
arXiv · 13 May 2026
ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows
arXiv · 13 May 2026
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
arXiv · 12 May 2026
How to cite this record
ethics.ai (13 May 2026), “Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy,” evidence record 4415, https://ethics.ai/record/4415 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.