Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization objectives can trade off or even conflict with each other
Record details
Published: 31 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 3 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
arXiv · 31 July 2026
Inference-Time Policy Alignment for Fair Reinforcement Learning
arXiv fairness query · 31 July 2026
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
arXiv · 30 July 2026
Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
arXiv cs.LG · 30 July 2026
RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
arXiv red teaming query · 29 July 2026
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
arXiv cs.LG · 29 July 2026
How to cite this record
ethics.ai (31 July 2026), “Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL,” evidence record 15675, https://ethics.ai/record/15675 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.