Weak-to-Strong On-Policy Distillation
On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prevailing approaches assume a teacher at least as capable as the student: they either distill a larger model into a smaller one, which fails at the frontier where no larger teacher exists, or consolidate multiple domain experts trained from a shared base, which requires costly training at the student's
Record details
Published: 28 July 2026
Source: arXiv cs.LG
Category: Research
Topics: Regulation · Children & education
Retrieved: 30 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Pass the Baton: Trajectory-Relayed On-Policy Distillation
arXiv cs.AI · 28 July 2026
Pictura: Perspective-View Self-Play at Scale for Driving
arXiv cs.AI · 28 July 2026
A Framework for Developing University Policies on Generative AI Governance: A Cross-national Comparative Study
arXiv cs.CY · 28 July 2026
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
HuggingFace Daily Papers · 29 July 2026
β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
HuggingFace Daily Papers · 29 July 2026
Flux-OPD: On-Policy Distillation with Evolving Contexts
HuggingFace Daily Papers · 29 July 2026
How to cite this record
ethics.ai (28 July 2026), “Weak-to-Strong On-Policy Distillation,” evidence record 14855, https://ethics.ai/record/14855 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.