A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether it mostly reshapes behavior the supervised model already has. Across a control grid of 18 runs that varies learning rate, KL weight, seed, initialization, and clipping, no configura
Record details
Published: 14 July 2026
Source: arXiv cs.AI
Category: Research
Topics: Regulation · Agents & autonomy
Retrieved: 15 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing
arXiv · 14 July 2026
UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies
arXiv cs.AI · 14 July 2026
Behavioral evolution and institutional coordination of multi-agent interactions in low-altitude airspace governance
Technological Forecasting and Social Change · 14 July 2026
From augmentation to autonomy: Artificial intelligence and the destabilization of agency in corporate governance
Futures (Elsevier) · 14 July 2026
Robots and post-retirement labor supply: Evidence from China
Telecommunications Policy · 14 July 2026
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
arXiv cs.AI · 14 July 2026
How to cite this record
ethics.ai (14 July 2026), “A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism,” evidence record 10466, https://ethics.ai/record/10466 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.