Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We study political censorship in Chinese-origin language models as a natural experiment, using probes, surgical ablations, and behavioral tests across nine open-weight models from five labs. Three findings follow. First, probe accuracy alone is non-diagnostic: politi
Record details
Published: 18 March 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
arXiv · 6 April 2026
Adoption and Effectiveness of AI-Based Anomaly Detection for Cross Provider Health Data Exchange
arXiv · 19 March 2026
Behavioural feasible set: Value alignment constraints on AI decision support
arXiv · 22 March 2026
Physiological and Semantic Patterns in Medical Teams Using an Intelligent Tutoring System
arXiv · 31 March 2026
AI Integrity: A New Paradigm for Verifiable AI Governance
arXiv · 13 April 2026
A Proposed Biomedical Data Policy Framework to Reduce Fragmentation, Improve Quality, and Incentivize Sharing in Indian Healthcare in the era of Artificial Intelligence and Digital Health
arXiv · 13 April 2026
How to cite this record
ethics.ai (18 March 2026), “Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails,” evidence record 7048, https://ethics.ai/record/7048 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.