The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically.This paper extends that challenge by examining two further possibilities: first, that evaluations of A
Record details
Published: 27 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
arXiv cs.CY · 30 July 2026
Evaluating whether AI models would sabotage AI safety research
arXiv · 27 April 2026
GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation
arXiv · 27 April 2026
Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
arXiv · 27 April 2026
Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture
arXiv · 26 April 2026
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
How to cite this record
ethics.ai (27 April 2026), “The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers,” evidence record 5324, https://ethics.ai/record/5324 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.