When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable. We study this problem in a hotel-pricing simulator where an agentic policy editor receives only region-level diagnostic feedback: summaries of how its price distribution differs from a benchmark policy across time, inventory, and market regions. The editor cannot observe benchmark actions, benchmark source code, rewar
Record details
Published: 3 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Healthcare · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Market-Alignment Risk in Pricing Agents: Trace Diagnostics and Trace-Prior RL under Hidden Competitor State
arXiv · 7 May 2026
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
arXiv · 28 May 2026
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
arXiv cs.CY · 14 August 2026
When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
arXiv · 18 May 2026
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
arXiv · 21 April 2026
How to cite this record
ethics.ai (3 July 2026), “When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions,” evidence record 273, https://ethics.ai/record/273 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.