DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD stil
Record details
Published: 6 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation · Children & education
Retrieved: 7 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI
arXiv · 6 August 2026
Implementation of Split Deadlines in a Large CS1 Course
arXiv · 7 August 2026
Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI
arXiv cs.CY · 7 August 2026
Teaching Intro AI When the Tools Can Do the Homework: A Course Redesign and a Student Bill of Rights
arXiv cs.CY · 7 August 2026
Building and Governing AI Systems: Advancing Social Workers' Roles across the Technology Industry, Human Service Organizations, and Policy Institutions
arXiv cs.CY · 6 August 2026
Where Does AI Innovation Go? Measuring Research Attention Imbalance in AI Music
arXiv · 7 August 2026
How to cite this record
ethics.ai (6 August 2026), “DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models,” evidence record 17336, https://ethics.ai/record/17336 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.