Auditing the Risk Claims of Distributional Reinforcement Learning
Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominan
Record details
Published: 13 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Safety & alignment · Agents & autonomy · Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm
arXiv · 11 July 2026
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
arXiv · 16 July 2026
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
arXiv · 16 July 2026
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
arXiv · 12 June 2026
Agent Economics: An Entropy-Controlled Pluralistic Alignment Framework for Preventing Artificial Hivemind in Autonomous Agents
arXiv · 8 June 2026
Gram: Assessing sabotage propensities via automated alignment auditing
arXiv · 28 May 2026
How to cite this record
ethics.ai (13 July 2026), “Auditing the Risk Claims of Distributional Reinforcement Learning,” evidence record 3038, https://ethics.ai/record/3038 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.