GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple turns of visual perception and action execution. However, existing methods still rely primarily on Supervised Fine-Tuning (SFT) with expert demonstrations, while the advanced reinforcement learning (RL) algorithm, specifically Group Relative Policy Optimization (GRPO), has not been effectively employed for multi-turn RL in these tasks because st
Record details
Published: 18 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Going Headless? On the Boundaries of Vertical AI Firms
arXiv · 18 May 2026
Ethical Hyper-Velocity (EHV): A Hardware-Rooted Zero-Trust Runtime Enforcement Architecture for Agentic AI Systems
arXiv · 18 May 2026
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
arXiv · 18 May 2026
REBAR: Reference Ethical Benchmark for Autonomy Readiness
arXiv · 18 May 2026
AI Agents May Always Fall for Prompt Injections
arXiv · 17 May 2026
When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
arXiv · 18 May 2026
How to cite this record
ethics.ai (18 May 2026), “GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents,” evidence record 4155, https://ethics.ai/record/4155 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.