How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate and amplifier are single heads; at larger scale they become bands of heads across adjacent layers. The gate contributes under 1% of output DLA, yet interchange testing (p = 120 detects the same motif in twelve models from six labs (2B to 72B), though specific
Record details
Published: 6 April 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
arXiv · 18 March 2026
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
arXiv · 5 April 2026
Incompleteness of AI Safety Verification via Kolmogorov Complexity
arXiv · 6 April 2026
Automated Analysis of Global AI Safety Initiatives: A Taxonomy-Driven LLM Approach
arXiv · 4 April 2026
Structural Rigidity and the 57-Token Predictive Window: A Physical Framework for Inference-Layer Governability in Large Language Models
arXiv · 4 April 2026
Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry
arXiv · 3 April 2026
How to cite this record
ethics.ai (6 April 2026), “How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models,” evidence record 6333, https://ethics.ai/record/6333 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.