Hidden Consensus:Preference-Validity Compression in Human Feedback
Standard RLHF pipelines often reduce heterogeneous human judgments into a single scalar reward target. We argue that this reduction can mis-measure alignment in structurally plural societies, where disagreement may reflect culturally, historically, linguistically, regionally, or normatively grounded interpretations rather than annotation noise. We call this failure Preference-Validity Compression, the collapse of multiple plural-valid response options into a single optimization target. Using Mal
Record details
Published: 9 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency
arXiv · 9 June 2026
Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-\texorpdfstring{$δ$}{delta} Alignment
arXiv · 9 June 2026
ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed Graphs
arXiv · 9 June 2026
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
arXiv · 9 June 2026
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
arXiv · 9 June 2026
Improving Text-Instance Alignment Of Foreground Conditioned Out-Painting Via Customized Concept Embedding
arXiv · 9 June 2026
How to cite this record
ethics.ai (9 June 2026), “Hidden Consensus:Preference-Validity Compression in Human Feedback,” evidence record 1214, https://ethics.ai/record/1214 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.