Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems
The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational risks. Most current evaluation frameworks neglect procedural compliance, leading to ''Machiavellian'' behaviors where agents strategically violate safety rules to maximize rewards - a direct manifestation of Goodhart's Law. To address this blind spot, we introduce MAC-Bench, a dynamic, adversarial benchmark designed to evaluate the procedural ali
Record details
Published: 5 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Why Memory Components Fail: Eight Years of License and Sustainability Events in Open-Source Data Infrastructure
arXiv · 5 June 2026
Blockchain Infrastructure for Intelligent Cyber--Physical--Social Systems:Post-Quantum Security, Interoperability, and Trustworthy Data Economies in the Era of Embodied AI
arXiv · 5 June 2026
Minimal Oversight: Uncertainty-Aware Governance for Delegated AI Systems
arXiv · 4 June 2026
TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
arXiv · 4 June 2026
Will the Agent Recuse Itself? Measuring LLM-Agent Compliance with In-Band Access-Deny Signals
arXiv · 4 June 2026
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
arXiv · 6 June 2026
How to cite this record
ethics.ai (5 June 2026), “Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems,” evidence record 1350, https://ethics.ai/record/1350 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.