When Does Muon Help Agentic Reinforcement Learning?
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The
Record details
Published: 16 July 2026
Source: HuggingFace Daily Papers
Category: Research
Topics: Regulation · Agents & autonomy
Retrieved: 21 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
RoboTTT: Context Scaling for Robot Policies
arXiv cs.AI · 16 July 2026
AutoSynthesis: An agentic system for automated meta-analysis
arXiv · 16 July 2026
Concept-Guided Spatial Regularization for World Models in Atari Pong
arXiv cs.AI · 16 July 2026
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
arXiv · 16 July 2026
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
arXiv · 16 July 2026
Traccia: An OpenTelemetry-Based Governance Platform for AI Systems
arXiv cs.CY · 17 July 2026
How to cite this record
ethics.ai (16 July 2026), “When Does Muon Help Agentic Reinforcement Learning?,” evidence record 11926, https://ethics.ai/record/11926 (originally published by HuggingFace Daily Papers).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.