Teach it to stop, not just to click
Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CUA) across five oracle-graded environments, we show a repaired policy's success rate is dominated by upstream variance: a variance-components decomposition across three cells (crossed data-draw $\times$ seed grid, bootstrap CIs) finds evaluation variance negligible ($σ_{\mathrm{eval}} \approx 0$) and the training-seed effect small everywhere ($\leq 10\%$); ins
Record details
Published: 19 July 2026
Source: arXiv cs.HC
Category: Research
Topics: Regulation · Agents & autonomy · Environment
Retrieved: 21 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Counterfactual Shapley Credit Assignment
arXiv cs.LG · 18 July 2026
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
arXiv · 17 July 2026
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
arXiv cs.CY · 16 July 2026
Multi-Turn On-Policy Distillation with Prefix Replay
HuggingFace Daily Papers · 15 July 2026
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
HuggingFace Daily Papers · 15 July 2026
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
arXiv · 23 July 2026
How to cite this record
ethics.ai (19 July 2026), “Teach it to stop, not just to click,” evidence record 12233, https://ethics.ai/record/12233 (originally published by arXiv cs.HC).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.