Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
Multi-Objective Alignment aims to align Large Language Models (LLMs) with diverse and often conflicting human values by optimizing multiple objectives simultaneously. Existing methods predominantly rely on static preference weight construction strategies. However, rigidly aligning to fixed targets discards valuable intermediate information, as training responses inherently embody valid preference trade-offs even when deviating from the target. To address this limitation, we propose Meal, i.e., M
Record details
Published: 27 April 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
arXiv · 18 May 2026
LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Last Week in AI · 21 July 2026
BayMOTH: Bayesian optiMizatiOn with meTa-lookahead -- a simple approacH
arXiv · 13 April 2026
Algorithmic Constitutionalism
arXiv · 16 May 2026
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
arXiv · 11 March 2026
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
arXiv · 9 March 2026
How to cite this record
ethics.ai (27 April 2026), “Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment,” evidence record 5320, https://ethics.ai/record/5320 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.