Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a
Record details
Published: 24 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Agents & autonomy
Retrieved: 27 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
arXiv cs.AI · 24 July 2026
Agentic Root Cause Analysis through Evidence-Grounded Reasoning
arXiv cs.AI · 24 July 2026
SceneActBench: Can Agents Act on the 3D Scenes They See?
arXiv · 24 July 2026
A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation
arXiv cs.AI · 24 July 2026
Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
arXiv · 24 July 2026
Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG
arXiv cs.AI · 24 July 2026
How to cite this record
ethics.ai (24 July 2026), “Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI,” evidence record 13699, https://ethics.ai/record/13699 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.