Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an outcome produced later by a particular target model, decoding policy, and judge. A filter tuned on a score that measures the wrong quantity spends its false positive budget on attacks
Record details
Published: 10 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Faithful or evasive? An empirical study on translation norm preferences of Chinese and American LLMs in Chinese official political and policy discourse
Frontiers in Artificial Intelligence · 10 August 2026
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
arXiv · 11 August 2026
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast
arXiv cs.LG · 9 August 2026
SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense
arXiv red teaming query · 11 August 2026
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
arXiv · 11 August 2026
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
arXiv red teaming query · 9 August 2026
How to cite this record
ethics.ai (10 August 2026), “Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks,” evidence record 18270, https://ethics.ai/record/18270 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.