One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
On a 300-persona life-simulation benchmark, pcsp achieves compositional zero-shot persona identification up to 17x above chance, Spearman rho approx 0.73 semantic-behavioral alignment, and 22x faster inference than an LLM-as-policy baseline. Life simulation games require hundreds to thousands of non-player characters (NPCs) that behave consistently with distinct personalities while remaining controllable through designer-authored natural language. Existing methods fail on constraints like person
Record details
Published: 22 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
arXiv · 22 May 2026
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
arXiv · 25 May 2026
REBAR: Reference Ethical Benchmark for Autonomy Readiness
arXiv · 18 May 2026
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
arXiv · 26 May 2026
Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems
arXiv · 26 May 2026
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
arXiv · 28 May 2026
How to cite this record
ethics.ai (22 May 2026), “One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents,” evidence record 3862, https://ethics.ai/record/3862 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.