How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
Alignment safety research assumes that ethical instructions improve model behavior, but how language models internally process such instructions remains unknown. We conducted over 600 multi-agent simulations across four models (Llama 3.3 70B, GPT-4o mini, Qwen3-Next-80B-A3B, Sonnet 4.5), four ethical instruction formats (none, minimal norm, reasoned norm, virtue framing), and two languages (Japanese, English). Confirmatory analysis fully replicated the Llama Japanese dissociation pattern from a
Record details
Published: 11 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness
arXiv · 3 August 2026
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
arXiv · 5 March 2026
Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems
arXiv · 17 March 2026
A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
arXiv · 25 May 2026
CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents
arXiv · 11 March 2026
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
arXiv · 9 March 2026
How to cite this record
ethics.ai (11 March 2026), “How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models,” evidence record 7400, https://ethics.ai/record/7400 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.