On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender ca
Record details
Published: 29 July 2026
Source: arXiv
Category: Research
Topics: Regulation
Retrieved: 30 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Anticipatory Data Governance in the Age of AI: Emerging Signals in Data Access, Reuse, and Sovereignty
arXiv · 29 July 2026
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
arXiv cs.AI · 29 July 2026
Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
arXiv cs.AI · 29 July 2026
RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
arXiv red teaming query · 29 July 2026
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
arXiv cs.LG · 29 July 2026
ReCo: Reweighting GRPO Against Distributional Concentration
arXiv cs.AI · 29 July 2026
How to cite this record
ethics.ai (29 July 2026), “On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment,” evidence record 14569, https://ethics.ai/record/14569 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.