Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones (0.91-0.98), while behavioral refusal drops to chance. This holds across four models and three families (1.5-3.8B, and at 14B). Refusal is therefore a shallow, response-site computation. We localize it to an early win
Record details
Published: 14 July 2026
Source: arXiv red teaming query
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 18 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
arXiv · 13 July 2026
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
arXiv · 15 July 2026
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
arXiv · 13 July 2026
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models
arXiv cs.LG · 13 July 2026
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
HuggingFace Daily Papers · 13 July 2026
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
arXiv · 15 July 2026
How to cite this record
ethics.ai (14 July 2026), “Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak,” evidence record 11622, https://ethics.ai/record/11622 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.