Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models. However, scaling OPD to long-horizon tasks exposes a critical flaw: as the student's generated prefix inevitably diverges from the teacher's thought process, the teacher's dense reward loses local exploitability. Continuing to generate and evaluate tokens on these ``drifted'' trajectories not only degrades reward quality but also incurs massive computational waste. To address this, we introduce \textbf{Prun
Record details
Published: 8 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Children & education
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
SOD: Step-wise On-policy Distillation for Small Language Model Agents
arXiv · 8 May 2026
Teachers' Perceived Benefits and Risks of AI Across Fifty-Five Countries: An Audit of LLM Alignment and Steerability
arXiv · 8 May 2026
Rubric-based On-policy Distillation
arXiv · 8 May 2026
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
arXiv · 10 May 2026
Age Verification in the Web -- Holy Grail to Control Access to Restricted Content
arXiv · 6 May 2026
Multi-Rollout On-Policy Distillation via Peer Successes and Failures
arXiv · 12 May 2026
How to cite this record
ethics.ai (8 May 2026), “Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning,” evidence record 4723, https://ethics.ai/record/4723 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.