Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six und
Record details
Published: 28 July 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy · Environment
Retrieved: 29 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
arXiv cs.AI · 28 July 2026
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
HuggingFace Daily Papers · 28 July 2026
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
HuggingFace Daily Papers · 28 July 2026
Constitutional governance for societies of AI agents in the built environment: a research agenda
arXiv cs.CY · 28 July 2026
Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
arXiv · 29 July 2026
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
arXiv · 29 July 2026
How to cite this record
ethics.ai (28 July 2026), “Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation,” evidence record 14133, https://ethics.ai/record/14133 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.