SWE-Touch tests coding agents in shared workspaces
The benchmark finds user counter-edits reduce coding-agent resolve rates on SWE-bench Verified.
Why it matters
The work highlights a gap between solo-agent coding benchmarks and real collaborative development, where code can change while an agent is working. It suggests coding agents need stronger awareness of evolving workspaces before they can be reliably used in shared repositories.
The key points
- 1.SWE-Touch targets user edits during coding-agent tasks.
- 2.Counter-Edits reduced average resolve rate by 7.7 points.
- 3.Failures were linked to limited evolving-workspace awareness.
Researchers introduced SWE-Touch, a framework for benchmarking coding agents when users inspect and modify code during an ongoing task. It injects validated Counter-Edits, described as plausible task-relevant code changes that conflict with task completion, along with contextual user messages. In tests of nine coding models, Counter-Edit lowered the average resolve rate by 7.7 percentage points on SWE-bench Verified, with degradation also reported on SWE-Bench Pro and DeepSWE.
⚡ Try this today
Use SWE-Touch-style tests to audit coding agents before relying on them in shared workspaces.
Sources & original reporting
This brief summarizes and links to reporting from the publishers below.
Enjoyed this brief? Get the next one in your inbox.
More in Research
New papers target VLMs' spatial reasoning gap
Researchers propose runtime memory, RL training, benchmarks and 3D generation methods for spatial AI.
HN readers debate practical limits of AI systems
Posts on drug discovery, math, privacy and AI labs drew discussion about where AI is useful and constrained.
New papers probe spatial intelligence in AI vision models
SpaRRTa, SMA and PinpointQA target spatial reasoning gaps in visual and multimodal systems.