SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对软件系统技术债务迁移问题,提出SWE Refactor Bench基准,通过三阶段评估协议衡量迁移完整性和行为正确性,测试编码代理执行能力。
📝 Abstract
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs ($5.4\%$) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores $47.0/100$. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, $58\%$ reach $99\%$ of the fixed checks, yet only $26\%$ reach $100\%$. Agent capability differs across migration categories: agents score $31.4$ on build toolchain rewrites but only $5.6$ on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
Problem

Research questions and friction points this paper is trying to address.

technical debt
whole-repository migration
coding agents
migration completeness
behavioural correctness
Innovation

Methods, ideas, or system contributions that make the work stand out.

SWE Refactor Bench
whole-repository migration
three-stage evaluation protocol
migration completeness
behavioural correctness
D
Deyao Hong
Navers Lab, Einsia.AI, Tsinghua University
Y
Yizhe Chi
Navers Lab, Einsia.AI, Tsinghua University
W
Wenyi Li
Navers Lab, Einsia.AI, Tsinghua University
X
Xiaoqiu Wang
Navers Lab, Einsia.AI, Tsinghua University
Mingju Gao
Mingju Gao
Unknown affiliation
Computer VisionRobotics
K
Kaisen Yang
Navers Lab, Einsia.AI, Tsinghua University
Bingxiang He
Bingxiang He
Second year PhD Candidate, Tsinghua University
Natural Language Processing
Y
Youjie Zheng
Navers Lab, Einsia.AI, Tsinghua University
C
Calvin Xiao
Navers Lab, Einsia.AI, Tsinghua University
Q
Qinhuai Na
Navers Lab, Einsia.AI, Tsinghua University