Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 in
Record details
Published: 16 July 2026
Source: HuggingFace Daily Papers
Category: Research
Topics: Agents & autonomy
Retrieved: 21 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
arXiv cs.AI · 16 July 2026
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
HuggingFace Daily Papers · 16 July 2026
When Does Muon Help Agentic Reinforcement Learning?
HuggingFace Daily Papers · 16 July 2026
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
HuggingFace Daily Papers · 16 July 2026
RecGPT-V3 Technical Report
HuggingFace Daily Papers · 16 July 2026
Recursive Harness Self-Improvement
HuggingFace Daily Papers · 16 July 2026
How to cite this record
ethics.ai (16 July 2026), “Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents,” evidence record 11920, https://ethics.ai/record/11920 (originally published by HuggingFace Daily Papers).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.