One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theo
Record details
Published: 12 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation · Agents & autonomy
Retrieved: 13 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
arXiv · 12 August 2026
Governing Agentic AI in FinTech
arXiv cs.CY · 13 August 2026
Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
arXiv · 12 August 2026
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
arXiv · 13 August 2026
Most biomedical publications show signs of LLM-assisted writing
arXiv cs.CY · 12 August 2026
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
arXiv · 11 August 2026
How to cite this record
ethics.ai (12 August 2026), “One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL,” evidence record 19026, https://ethics.ai/record/19026 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.