The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
RLHF-aligned language models exhibit response homogenization: on TruthfulQA (n=790), 40-79% of questions produce a single semantic cluster across 10 i.i.d. samples. On affected questions, sampling-based uncertainty methods have zero discriminative power (AUROC=0.500), while free token entropy retains signal (0.603). This alignment tax is task-dependent: on GSM8K (n=500), token entropy achieves 0.724 (Cohen's d=0.81). A base-vs-instruct ablation confirms the causal role of alignment: the base mod
Record details
Published: 25 March 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Latent Bias Alignment for High-Fidelity Diffusion Inversion in Real-World Image Reconstruction and Manipulation
arXiv · 25 March 2026
BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
arXiv · 25 March 2026
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
arXiv · 26 March 2026
On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers
arXiv · 30 March 2026
Aligning Recommendations with User Popularity Preferences
arXiv · 1 April 2026
Prompt Programming for Cultural Bias and Alignment of Large Language Models
arXiv · 17 March 2026
How to cite this record
ethics.ai (25 March 2026), “The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation,” evidence record 6772, https://ethics.ai/record/6772 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.