The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
On-policy distillation (OPD) and on-policy self-distillation (OPSD) have emerged as promising post-training methods for large language models, offering dense token-level supervision on trajectories sampled from the model's own policy. However, existing results on their effectiveness remain mixed: while OP(S)D has shown promise in system prompt and knowledge internalization, recent studies also report instability and degradation. In this work, we present a comprehensive empirical study of when OP
Record details
Published: 11 May 2026
Source: arXiv
Category: Research
Topics: Regulation
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
arXiv · 11 May 2026
Switching-Geometry Analysis of Deflated Q-Value Iteration
arXiv · 11 May 2026
PhyGround: Benchmarking Physical Reasoning in Generative World Models
arXiv · 11 May 2026
Polyendocrine metabolic ovarian syndrome, the new name for polycystic ovary syndrome: a multistep global consensus process
OpenAlex · 12 May 2026
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
arXiv · 12 May 2026
LLM-X: A Scalable Negotiation-Oriented Exchange for Communication Among Personal LLM Agents
arXiv · 12 May 2026
How to cite this record
ethics.ai (11 May 2026), “The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes,” evidence record 4534, https://ethics.ai/record/4534 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.