Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tension between competing goals dictates that aggressively optimizing for one metric (e.g., helpfulness) frequently incurs a substantial penalty on another (e.g., harmlessness). While prior work mainly focuses on data selection, parameter merging, or algorithmic balancing during training, these approaches merely force compr
Record details
Published: 12 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arXiv · 12 May 2026
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
arXiv · 13 May 2026
STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
arXiv · 13 May 2026
CAMAL: Improving Attention Alignment and Faithfulness with Segmentation Masks
arXiv · 8 May 2026
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
arXiv · 8 May 2026
Strat-LLM: Stratified Strategy Alignment for LLM-based Stock Trading with Real-time Multi-Source Signals
arXiv · 7 May 2026
How to cite this record
ethics.ai (12 May 2026), “Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion,” evidence record 4498, https://ethics.ai/record/4498 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.