ReCo: Reweighting GRPO Against Distributional Concentration
Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We trace this concentration to two mechanisms in the GRPO upd
Record details
Published: 29 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation
Retrieved: 30 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
arXiv cs.AI · 29 July 2026
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
arXiv cs.LG · 29 July 2026
Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
arXiv cs.AI · 29 July 2026
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
arXiv cs.AI · 29 July 2026
Anticipatory Data Governance in the Age of AI: Emerging Signals in Data Access, Reuse, and Sovereignty
arXiv · 29 July 2026
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
arXiv · 29 July 2026
How to cite this record
ethics.ai (29 July 2026), “ReCo: Reweighting GRPO Against Distributional Concentration,” evidence record 14819, https://ethics.ai/record/14819 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.