AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, standard algorithms like GRPO apply sequence-level rewards uniformly to all tokens, creating a severe credit-assignment bottleneck. While on-policy self-distillation attempts to resolve this by conditioning a self-teacher on privileged contexts, direct exposure to raw oracle solutions often induces over-conditioned teacher distributions, implicit a
Record details
Published: 18 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
arXiv · 27 April 2026
LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Last Week in AI · 21 July 2026
Algorithmic Constitutionalism
arXiv · 16 May 2026
BayMOTH: Bayesian optiMizatiOn with meTa-lookahead -- a simple approacH
arXiv · 13 April 2026
Meta-Transfer Learning for mmWave Beam Alignment
arXiv · 1 July 2026
A systematic review of toxicity in large language models: definitions, datasets, detectors, detoxification methods and challenges
Artificial Intelligence Review · 2 July 2026
How to cite this record
ethics.ai (18 May 2026), “AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment,” evidence record 4110, https://ethics.ai/record/4110 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.