Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models
Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly. In practice, the computational expense of different attack strategies can vary by orders of magnitude. Consequently, ASR at a fixed budget can obscure the true effort required to jailbreak a model, thereby making it hard to determine whether an attack's cost justifies its payoff to the attacker. We propose a co
Record details
Published: 9 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Role of Feedback Alignment in Self-Distillation
arXiv · 9 June 2026
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
arXiv · 9 June 2026
FreeBridge: Variational Schrödinger Bridges for Cellular Transition Dynamics
arXiv · 9 June 2026
Improving Text-Instance Alignment Of Foreground Conditioned Out-Painting Via Customized Concept Embedding
arXiv · 9 June 2026
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
arXiv · 9 June 2026
Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$δ$}{delta} Alignment
arXiv · 9 June 2026
How to cite this record
ethics.ai (9 June 2026), “Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models,” evidence record 1178, https://ethics.ai/record/1178 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.