Invariant Gradient Alignment for Robust Reasoning Distillation
Large language models (LLMs) suffer from shortcut learning: they systematically fail on out-of-distribution (OOD) inputs whose semantic surface differs from training data, even when the logical structure is identical. This undermines knowledge distillation pipelines that transfer chain-of-thought reasoning to smaller students. We introduce Invariant Gradient Alignment (IGA), a training framework that aligns gradient updates across semantically diverse but logically isomorphic examples via three
Record details
Published: 3 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Children & education
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises
arXiv · 1 June 2026
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
arXiv · 9 June 2026
Rationalize: Shared Semantic Reasoning for Human-AI Alignment
arXiv · 28 May 2026
The Role of Feedback Alignment in Self-Distillation
arXiv · 9 June 2026
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
arXiv · 24 May 2026
Improving Capstone Team Outcomes through Dynamic Skill Matching and Preference Alignment
arXiv · 14 June 2026
How to cite this record
ethics.ai (3 June 2026), “Invariant Gradient Alignment for Robust Reasoning Distillation,” evidence record 1489, https://ethics.ai/record/1489 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.