Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positiv
Record details
Published: 30 July 2026
Source: arXiv red teaming query
Category: Research
Topics: Agents & autonomy · Transparency
Retrieved: 3 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
"Nobody Did This": Contribution, Originality, and Accountability in Agent-Mediated Collaboration
arXiv cs.CY · 30 July 2026
One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
arXiv cs.AI · 30 July 2026
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
arXiv cs.AI · 29 July 2026
Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM
arXiv cs.AI · 29 July 2026
Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm
arXiv · 28 July 2026
Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels
arXiv cs.CY · 28 July 2026
How to cite this record
ethics.ai (30 July 2026), “Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks,” evidence record 15793, https://ethics.ai/record/15793 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.