A Revealed Preference Framework for AI Alignment
Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal's preferences or pursue its own? To study this question using revealed preference techniques, I introduce the Luce Alignment Model, where the AI's choices are a mixture of two Luce rules, one reflecting the human's preferences and the other the AI's. I show that the AI's alignment (similarity of human and AI preferences) can be generically identified in two sett
Record details
Published: 29 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Reward Hacking as Equilibrium under Finite Evaluation
arXiv · 30 March 2026
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
arXiv · 31 March 2026
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
arXiv · 27 March 2026
Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles
arXiv · 27 March 2026
The Persistent Vulnerability of Aligned AI Systems
arXiv · 31 March 2026
From Logic Monopoly to Social Contract: Separation of Power and the Institutional Foundations for Autonomous Agent Economies
arXiv · 26 March 2026
How to cite this record
ethics.ai (29 March 2026), “A Revealed Preference Framework for AI Alignment,” evidence record 6612, https://ethics.ai/record/6612 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.