GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likelihood. A dominant and efficient family of methods replaces the likelihood in standard RL with its evidence lower bound (ELBO), estimated from randomly masked sequences. Despite being well aligned with pre-training, these approaches introduce bias through training--inference mismatch by using the ELBO as a likelihood sur
Record details
Published: 28 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Regulation
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis
arXiv · 29 May 2026
Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions -- A Governance Framework for High-Stakes AI Systems
arXiv · 27 May 2026
A Policy-Driven Runtime Layer for Agentic LLM Serving
arXiv · 26 May 2026
COPF: An Online Framework for Deployment-Stable Counterfactual Fairness in Evolving Graphs
arXiv · 30 May 2026
Silent Failures in Federated Personalization of Foundation Models
arXiv · 31 May 2026
AGORA: Can Deliberation and Governance Gates Absorb Participation Bias in Transit Planning?
arXiv · 31 May 2026
How to cite this record
ethics.ai (28 May 2026), “GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models,” evidence record 3538, https://ethics.ai/record/3538 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.