The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When no scratchpad is visible, behavior alone cannot distinguish strategic from genuine compliance. We ask whether hidden states reveal what outputs hide. We run a 13-model sweep for naturally-emerging faking, then probe and steer hidden states on the two models that fake. Natural faking appears only in Qwen3-32B (+18.2pp) and Llama-3.1-8B (+24.4pp at n=
Record details
Published: 15 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 16 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
arXiv red teaming query · 14 July 2026
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
arXiv · 15 July 2026
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
arXiv · 13 July 2026
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
arXiv · 13 July 2026
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models
arXiv cs.LG · 13 July 2026
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
arXiv · 16 July 2026
How to cite this record
ethics.ai (15 July 2026), “The Refusal Residue: When Probes Catch Alignment Faking and When They Don't,” evidence record 10606, https://ethics.ai/record/10606 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.