A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated report. We ask what this does to a class of defect no single worker can see: a contradiction in the relation between two distant sections of a document. Holding the documents, defects, mechanism, scoring, and seed fixed, we vary only the model -- ten systems across five generations from one developer and five providers from distinct alignment para
Record details
Published: 25 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems
arXiv · 17 March 2026
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
arXiv · 11 March 2026
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
arXiv · 5 March 2026
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
arXiv · 25 May 2026
EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
arXiv · 26 May 2026
Position: AI Safety Requires Effective Controllability
arXiv · 26 May 2026
How to cite this record
ethics.ai (25 May 2026), “A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration,” evidence record 3756, https://ethics.ai/record/3756 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.