How LLMs Are Persuaded: A Few Attention Heads, Rerouted
Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly understood. We uncover a compact causal mechanism for persuasion-induced factual errors. A small set of mid-layer attention heads almost entirely determines the model's answer. These heads write answer options into a low-dimensional polyhedron, with options occupying distinct vertices. Persuasion does not blur belief or merely reduce confidence; it
Record details
Published: 10 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step
arXiv · 10 May 2026
The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?
arXiv · 10 May 2026
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
arXiv · 10 May 2026
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
arXiv · 9 May 2026
Open Ontologies: Tool-Augmented Ontology Engineering with Stable Matching Alignment
arXiv · 9 May 2026
Data-driven Circuit Discovery for Interpretability of Language Models
arXiv · 9 May 2026
How to cite this record
ethics.ai (10 May 2026), “How LLMs Are Persuaded: A Few Attention Heads, Rerouted,” evidence record 4635, https://ethics.ai/record/4635 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.