Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 in
Record details
Published: 16 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Agents & autonomy
Retrieved: 18 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
HuggingFace Daily Papers · 16 July 2026
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
arXiv cs.AI · 16 July 2026
RoboTTT: Context Scaling for Robot Policies
arXiv cs.AI · 16 July 2026
An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions
JMIR (Journal of Medical Internet Research) · 16 July 2026
AutoSynthesis: An agentic system for automated meta-analysis
arXiv · 16 July 2026
When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
arXiv cs.AI · 16 July 2026
How to cite this record
ethics.ai (16 July 2026), “Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents,” evidence record 11588, https://ethics.ai/record/11588 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.