Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics
Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a 2$\times$2 taxonomy -- epistemic versus motivational reason
Record details
Published: 10 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Healthcare · Agents & autonomy
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations
arXiv cs.LG · 10 August 2026
SHRIMP: Iterative Refinement of Robot Task Plans
arXiv cs.HC · 9 August 2026
MIRA: Medical Image Reflection for Agentic Diagnosis
arXiv cs.AI · 11 August 2026
SocialFiVis: A Visual Analytics Sandbox for LLM-Grounded Multi-Agent Simulation in Social Finance
arXiv cs.HC · 9 August 2026
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
arXiv cs.AI · 12 August 2026
Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits
arXiv cs.CY · 13 August 2026
How to cite this record
ethics.ai (10 August 2026), “Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics,” evidence record 18267, https://ethics.ai/record/18267 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.