I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization
Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this gap with dense token supervision, yet applying it throughout training creates a different failure mode: the teacher is a biased, low-variance surrogate for the reward objective, so persistent imitation can oppose reward-improving updates after the policy becomes capable of produ
Record details
Published: 13 August 2026
Source: arXiv cs.LG
Category: Research
Topics: Bias & fairness · Regulation
Retrieved: 14 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Variable Selection in the Context of AI Fairness
arXiv cs.CY · 13 August 2026
Language Models as a Challenge for Business Ethics – A Partially Open-Source Approach
Science and Engineering Ethics · 14 August 2026
IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning
arXiv cs.LG · 11 August 2026
Manipulation-Proof Oblivious Audits against Deceptive Model Providers
arXiv · 5 August 2026
Variable Selection in the Context of AI Fairness
arXiv fairness query · 4 August 2026
Progressive in Principle, Centrist in Practice: LLM Political Bias Is Instrument-Dependent
arXiv cs.CY · 4 August 2026
How to cite this record
ethics.ai (13 August 2026), “I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization,” evidence record 19475, https://ethics.ai/record/19475 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.