SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning. We propose SEED (SElf-Evolving On-Policy Distillation), a self-evolving
Record details
Published: 15 July 2026
Source: HuggingFace Daily Papers
Category: Research
Topics: Regulation · Agents & autonomy · Environment
Retrieved: 17 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Multi-Turn On-Policy Distillation with Prefix Replay
HuggingFace Daily Papers · 15 July 2026
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
arXiv cs.CY · 16 July 2026
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
arXiv · 14 July 2026
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
arXiv · 17 July 2026
Counterfactual Shapley Credit Assignment
arXiv cs.LG · 18 July 2026
Teach it to stop, not just to click
arXiv cs.HC · 19 July 2026
How to cite this record
ethics.ai (15 July 2026), “SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning,” evidence record 10995, https://ethics.ai/record/10995 (originally published by HuggingFace Daily Papers).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.