Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning
In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization. To address this issue, we develop a behavioral advantage corrected policy evaluation (BAC-PE) approach, which utilizes the \emph{Q}-function of the behavior policy to correct the learned policy's \emph{Q}-function, thus mitigating pessimistic conservatism and overestimation bias. Fur
Record details
Published: 3 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Bias & fairness · Regulation
Retrieved: 4 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints
arXiv fairness query · 3 August 2026
AIDLC–governance indicator framework: a lifecycle-based approach to institutional AI governance
AI & Society · 4 August 2026
Whose ethics? Whose AI? Refining the Philippine AI Regulation Act toward a Contextual AI Ethic
AI & Society · 4 August 2026
Progressive in Principle, Centrist in Practice: LLM Political Bias Is Instrument-Dependent
arXiv cs.CY · 4 August 2026
Variable Selection in the Context of AI Fairness
arXiv fairness query · 4 August 2026
Manipulation-Proof Oblivious Audits against Deceptive Model Providers
arXiv · 5 August 2026
How to cite this record
ethics.ai (3 August 2026), “Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning,” evidence record 16106, https://ethics.ai/record/16106 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.