Fragility of Value under Imperfect Alignment
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optim
Record details
Published: 30 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 3 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
arXiv · 31 July 2026
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
arXiv · 31 July 2026
The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning
arXiv cs.CY · 30 July 2026
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
arXiv cs.CY · 30 July 2026
Inference-Time Policy Alignment for Fair Reinforcement Learning
arXiv fairness query · 31 July 2026
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
HuggingFace Daily Papers · 31 July 2026
How to cite this record
ethics.ai (30 July 2026), “Fragility of Value under Imperfect Alignment,” evidence record 15687, https://ethics.ai/record/15687 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.