SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to messages. This leads us to ask: how do coding agents understand and respond to code changes in a shared workspace? We introduce SWE-Touch, a framework that stress-tests this setting through validated Counter-Edits: plausible edits to
Record details
Published: 3 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Agents & autonomy
Retrieved: 4 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation
arXiv cs.AI · 3 August 2026
Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation
arXiv cs.AI · 3 August 2026
Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment
arXiv cs.AI · 3 August 2026
Real-Time Detection and Repair of LLM Agent Failures
arXiv cs.AI · 3 August 2026
A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI
arXiv cs.AI · 3 August 2026
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision
arXiv cs.AI · 3 August 2026
How to cite this record
ethics.ai (3 August 2026), “SWE-Touch: Benchmarking Coding Agents When Users Touch the Code,” evidence record 16091, https://ethics.ai/record/16091 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.