Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial perturbations alter the model's internal reasoning. Consequently, the mechanisms underlying unsafe or incorrect behaviors remain poorly understood. We introduce a mechanistic framework for diagnosing LLM vulner
Record details
Published: 8 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Alignment Plausibility: A New Standard for Assuring AI in Healthcare
arXiv · 8 July 2026
Interpretable polyp classification via end-to-end Concept Bottleneck Models with vision-language concept alignment
Frontiers in Artificial Intelligence · 10 July 2026
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
arXiv · 7 July 2026
Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations
arXiv · 6 July 2026
Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health
arXiv · 12 July 2026
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
arXiv · 3 July 2026
How to cite this record
ethics.ai (8 July 2026), “Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs,” evidence record 110, https://ethics.ai/record/110 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.