APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
Aligning large language models (LLMs) with diverse human preferences requires pluralistic alignment, where a single model must respect the values of multiple distinct groups simultaneously. In federated reinforcement learning from human feedback (FedRLHF), these groups align a shared policy without centralizing preference data, which makes fair reward aggregation essential. Existing aggregation methods exhibit clear trade offs: average based aggregation systematically under aligns worst performi
Record details
Published: 5 April 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
arXiv · 6 April 2026
Incompleteness of AI Safety Verification via Kolmogorov Complexity
arXiv · 6 April 2026
Automated Analysis of Global AI Safety Initiatives: A Taxonomy-Driven LLM Approach
arXiv · 4 April 2026
Structural Rigidity and the 57-Token Predictive Window: A Physical Framework for Inference-Layer Governability in Large Language Models
arXiv · 4 April 2026
Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry
arXiv · 3 April 2026
From Reactive to Proactive: A Multi-Regulatory Empirical Analysis of 480 AI Incidents and a Data-Driven Governance Compliance Framework
arXiv · 10 April 2026
How to cite this record
ethics.ai (5 April 2026), “APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs,” evidence record 6337, https://ethics.ai/record/6337 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.