Constitutional On-Policy Safe Distillation
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervision. Prior work has shown that OPSD can collapse in verifiable reasoning tasks, but safety alignment differs in that it is guided by high-level constitutions rather than explicit target answers, making it a natural setting to revisit dense distillation. However, our pilot study show that safety OPSD still suffers from
Record details
Published: 2 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Beyond Static Priors: Dynamic Neural Guidance for Large-Scale Ant Colony Optimization
arXiv · 2 June 2026
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
arXiv · 1 June 2026
AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task
arXiv · 2 June 2026
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization
arXiv · 3 June 2026
Silent Failures in Federated Personalization of Foundation Models
arXiv · 31 May 2026
When AI Says It Feels
arXiv · 4 June 2026
How to cite this record
ethics.ai (2 June 2026), “Constitutional On-Policy Safe Distillation,” evidence record 3247, https://ethics.ai/record/3247 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.