Misaligned by Reward: Socially Undesirable Preferences in LLMs
Reward models are a key component of large language model alignment, serving as proxies for human preferences during training. However, existing evaluations focus primarily on broad instruction-following benchmarks, providing limited insight into whether these models capture socially desirable preferences. As a result, important failures in social alignment can remain hidden. We extend reward-model benchmarking to four socially consequential domains: bias, safety, morality, and ethical reasoning
Record details
Published: 6 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
arXiv · 6 May 2026
Resolving the bias-precision paradox with stochastic causal representation learning for personalized medicine
arXiv · 7 May 2026
Strat-LLM: Stratified Strategy Alignment for LLM-based Stock Trading with Real-time Multi-Source Signals
arXiv · 7 May 2026
Brainrot: Deskilling and Addiction are Overlooked AI Risks
arXiv · 5 May 2026
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
arXiv · 4 May 2026
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
arXiv · 8 May 2026
How to cite this record
ethics.ai (6 May 2026), “Misaligned by Reward: Socially Undesirable Preferences in LLMs,” evidence record 4904, https://ethics.ai/record/4904 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.