Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but often off-policy masked states, while reinforcement learning (RL) relies on sparse rewards or value modeling. This paper proposes \textbf{trace-based on-policy distillation (TOPD)}, a teacher-supervised framework that transfers reasoning ability to a target dLLM without
Record details
Published: 18 July 2026
Source: arXiv cs.LG
Category: Research
Topics: Regulation
Retrieved: 21 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Distilled Reinforcement Learning for LLM Post-training
HuggingFace Daily Papers · 18 July 2026
Migration and Displacement in Crisis: Protection, Policy, and Practice
Harvard Kennedy School · 18 July 2026
Counterfactual Shapley Credit Assignment
arXiv cs.LG · 18 July 2026
The shared blind spot: why diverse AI governance approaches fail for the same reason
AI & Society · 19 July 2026
A comparative and integrative mapping of thematic trajectories, the digital turn, and future research frontiers in telecommunications policy
Telecommunications Policy · 19 July 2026
Working remotely, unequally: Evidence from a digitally divided labor market
Telecommunications Policy · 19 July 2026
How to cite this record
ethics.ai (18 July 2026), “Trace-Based On-Policy Distillation for Masked Diffusion Language Models,” evidence record 12272, https://ethics.ai/record/12272 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.