Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly understood. This paper provides a systematic investigation of OPD dynamics and mechanisms. We first identify that two conditions govern whether OPD succeeds or fails: (i) the student and teacher should share compatible thinking patterns; and (ii) even with consistent thinking patterns and higher scores, the teacher must offer genuinely new capabilities b
Record details
Published: 14 April 2026
Source: arXiv
Category: Research
Topics: Regulation · Children & education · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Enabling and Inhibitory Pathways of Students' AI Use Concealment Intention in Higher Education: Evidence from SEM and fsQCA
arXiv · 13 April 2026
A Systematic AI Adoption Framework for Higher Education: From Student GenAI Usage to Institutional Integration
arXiv · 23 April 2026
How unique are hallucinated citations offered by generative Artificial Intelligence models?
arXiv · 31 March 2026
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
arXiv · 10 May 2026
Knowing the Rules Is Not Enough: Student Regulatory Awareness and Use of GenAI in Higher Education
arXiv · 16 May 2026
Interpretable Markov-Based Spatiotemporal Risk Surfaces for Missing-Child Search Planning with Reinforcement Learning and LLM-Based Quality Assurance
arXiv · 9 March 2026
How to cite this record
ethics.ai (14 April 2026), “Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe,” evidence record 5860, https://ethics.ai/record/5860 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.