Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not how it unfolded. Two attacks that produce equally harmful outputs may have followed completely different paths, and ASR cannot tell them apart. We make those hidden paths observable from logits alone. Temporal Logit Observability (TLO) is a training-free diagnostic that watches a compliance-refusal margin during decoding and places each model-attac
Record details
Published: 28 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Healthcare
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
arXiv · 28 May 2026
SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models
arXiv · 27 May 2026
Next-Billion AI Index: The compass for AI utility and adoption in the global majority
arXiv · 29 May 2026
Unsupervised Pattern Analysis in Japanese Veterinary Toxicology: A Regulatory-Compliant Framework for Cross-Species Risk Assessment
arXiv · 4 June 2026
The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In
arXiv · 6 June 2026
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
How to cite this record
ethics.ai (28 May 2026), “Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures,” evidence record 3520, https://ethics.ai/record/3520 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.