GTA: Generating Long-Horizon Tasks for Web Agents at Scale
Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limited by the lack of scalable, process-level supervision. Existing benchmarks are largely manually constructed, providing only coarse start-goal annotations without intermediate trajectories, while recent automatic generation efforts remain expensive, biased, and shallow. These limitations prevent reliable training and evaluation of agents that mus
Record details
Published: 28 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Entity-Collision: A Stratified Protocol for Attributing Retrieval Lift in Agent Memory
arXiv · 28 May 2026
Examining Agents' Bias Amplification versus Suppression in Multi-Agent Systems
arXiv · 27 May 2026
A Policy-Driven Runtime Layer for Agentic LLM Serving
arXiv · 26 May 2026
Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation
arXiv · 30 May 2026
AGORA: Can Deliberation and Governance Gates Absorb Participation Bias in Transit Planning?
arXiv · 31 May 2026
Symbolic Reasoning Frameworks Trigger Memory-Mediated Ecosystem Dynamics in Multi-Agent LLM Systems
arXiv · 22 May 2026
How to cite this record
ethics.ai (28 May 2026), “GTA: Generating Long-Horizon Tasks for Web Agents at Scale,” evidence record 3551, https://ethics.ai/record/3551 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.