What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite safety training. While most defenses operate at the prompt or output level, it remains unclear how harmful intent is encoded within the model's internal representations. We investigate this question by analyzing token-level predictive entropy trajectories across layers of a frozen LLM using the logit lens. We find that static aggregate statistic
Record details
Published: 23 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs
arXiv · 21 June 2026
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
arXiv · 15 July 2026
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
HuggingFace Daily Papers · 26 July 2026
TRuE-XAI: causal and explainable ai framework for trustworthy corporate earnings growth forecasting
Frontiers in Artificial Intelligence · 27 July 2026
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
arXiv · 27 July 2026
Physiological and Semantic Patterns in Medical Teams Using an Intelligent Tutoring System
arXiv · 31 March 2026
How to cite this record
ethics.ai (23 June 2026), “What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics,” evidence record 628, https://ethics.ai/record/628 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.