Counsel: A Meta-Evaluation Dataset for Agentic Tasks
As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agentic benchmarks can take hours, making it difficult to scale evaluations for measuring performance or curating training data. This has driven widespread reliance on automated approaches such as LLM-as-a-judge (LLMJ) to critique agents at the process and outcome-levels at scale, however, the soundness of LLMJ critiques o
Record details
Published: 19 June 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Knowledge Reutilization in Meta-Reinforcement Learning
arXiv · 16 June 2026
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
arXiv · 3 June 2026
STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
arXiv · 19 May 2026
SkillEvolver: Skill Learning as a Meta-Skill
arXiv · 11 May 2026
Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness
arXiv · 3 August 2026
MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph
arXiv · 11 August 2026
How to cite this record
ethics.ai (19 June 2026), “Counsel: A Meta-Evaluation Dataset for Agentic Tasks,” evidence record 758, https://ethics.ai/record/758 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.