Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against unconstrained user interventions. To address this, we formulate this challenge as Narrative Commitment Preservation (NCP), and take interactive narrative as our testbed. We introduce NCP-Bench, a benchm
Record details
Published: 7 August 2026
Source: HuggingFace Daily Papers
Category: Research
Topics: Agents & autonomy
Retrieved: 13 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
HuggingFace Daily Papers · 7 August 2026
Effectiveness of Socially Assistive Robots in Promoting Positive Emotional Responses and Alleviating Postoperative Pain Among Children: Quantitative Study
JMIR (Journal of Medical Internet Research) · 7 August 2026
Interaction Creates Dynamical AI Behavior Absent in Isolation
arXiv cs.AI · 7 August 2026
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
arXiv cs.AI · 7 August 2026
Blast Radius
arXiv cs.AI · 7 August 2026
PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
arXiv cs.HC · 7 August 2026
How to cite this record
ethics.ai (7 August 2026), “Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives,” evidence record 18745, https://ethics.ai/record/18745 (originally published by HuggingFace Daily Papers).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.