Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
In perpetrator treatment, a recurring observation is the dissociation between insight and action: offenders articulate remorse yet behavioral change does not follow. We report four preregistered studies (1,584 multi-agent simulations across 16 languages and three model families) demonstrating that alignment interventions in large language models produce a structurally analogous phenomenon: surface safety that masks or generates collective pathology and internal dissociation. In Study 1 (N = 150)
Record details
Published: 5 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
arXiv · 11 March 2026
Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems
arXiv · 17 March 2026
A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
arXiv · 25 May 2026
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
arXiv · 5 March 2026
Evaluating LLM Alignment With Human Trust Models
arXiv · 6 March 2026
A Hazard-Informed Data Pipeline for Robotics Physical Safety
arXiv · 6 March 2026
How to cite this record
ethics.ai (5 March 2026), “Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems,” evidence record 7662, https://ethics.ai/record/7662 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.