REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation, but apply a single global coefficient $λ$ to every token. This can drive the student to fit extreme peaks in the implicit reward, causing reward hacking and unstable training, and the optimal $λ$ varies across domains, requiring costly sweeps. We prop
Record details
Published: 12 August 2026
Source: arXiv cs.LG
Category: Research
Topics: Regulation · Children & education
Retrieved: 14 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
arXiv · 12 August 2026
Trump wants MMR vaccine split up: the science behind why it’s a bad idea
Nature Machine Intelligence · 11 August 2026
Social Media as a Driver of Obesity in Children and Adolescents (Aged 6-18 Years): It Is Time for Regulatory Action
JMIR (Journal of Medical Internet Research) · 10 August 2026
Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension
arXiv · 13 August 2026
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
arXiv cs.AI · 10 August 2026
Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation
arXiv cs.AI · 10 August 2026
How to cite this record
ethics.ai (12 August 2026), “REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation,” evidence record 19481, https://ethics.ai/record/19481 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.