ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization
Offline-to-online reinforcement learning harnesses the stability of offline pretraining and the flexibility of online fine-tuning. A key challenge lies in the non-stationary distribution shift between offline datasets and the evolving online policy. Common approaches often rely on static mixing ratios or heuristic-based replay strategies, which lack adaptability to different environments and varying training dynamics, resulting in suboptimal tradeoff between stability and asymptotic performance.
Record details
Published: 14 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
GAGPO: Generalized Advantage Grouped Policy Optimization
arXiv · 13 May 2026
Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning
arXiv · 13 May 2026
LLM-X: A Scalable Negotiation-Oriented Exchange for Communication Among Personal LLM Agents
arXiv · 12 May 2026
When Does Non-Uniform Replay Matter in Reinforcement Learning?
arXiv · 11 May 2026
Jurisdiction over Ubiquitous Copyright Infringements: Should Right-Holders Be Allowed to Sue at Home?
arXiv · 17 May 2026
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
arXiv · 18 May 2026
How to cite this record
ethics.ai (14 May 2026), “ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization,” evidence record 4340, https://ethics.ai/record/4340 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.