Inference-Time Policy Alignment for Fair Reinforcement Learning
Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. For instance, an agent trained to maximize expected cumulative reward may not accommodate previously unknown stakeholder preferences. Existing approaches to achieve fairness, a type of preference, in RL typically assume that such preferences are known a priori and require
Record details
Published: 31 July 2026
Source: arXiv fairness query
Category: Research
Topics: Bias & fairness · Regulation · Safety & alignment · Agents & autonomy
Retrieved: 4 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
LLM Constitutional Multi-Agent Governance
arXiv · 13 March 2026
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
arXiv · 31 July 2026
Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints
arXiv fairness query · 3 August 2026
Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI
arXiv cs.CY · 28 July 2026
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
arXiv · 27 July 2026
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
HuggingFace Daily Papers · 4 August 2026
How to cite this record
ethics.ai (31 July 2026), “Inference-Time Policy Alignment for Fair Reinforcement Learning,” evidence record 16134, https://ethics.ai/record/16134 (originally published by arXiv fairness query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.