ChatPaper.aiChatPaper

SWE Refactor Bench: 코딩 에이전트가 장기적 전체 저장소 스택 마이그레이션을 완료할 수 있는가?

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

August 24, 2026
저자: Deyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na
cs.AI

초록

현대 소프트웨어 시스템은 수십 년간의 개발 과정에서 기술 부채(technical debt)를 축적하며, 이로 인해 마이그레이션은 비용이 많이 들고 대부분 수작업으로 수행된다. 버그 수정에 점점 더 능숙해지는 코딩 에이전트가 이러한 마이그레이션을 자율적으로 수행할 수 있을까? 기존 벤치마크는 이 질문에 답할 수 없다. 벤치마크가 동작의 정확성만 평가할 뿐 마이그레이션이 실제로 발생했는지 여부는 평가하지 않기 때문이다. 이로 인해 에이전트가 테스트를 통과시키기 위해 원래 구현을 복사하는 쉬운 우회 방법이 가능하며, 우리는 이를 맹점(Blindness)이라고 부른다. 이 문제를 해결하기 위해 우리는 SWE Refactor Bench를 소개한다. 이는 4가지 유형의 기술 부채를 포괄하는 20개의 전체 저장소 마이그레이션으로 구성된 벤치마크이다. 3단계 평가 프로토콜은 마이그레이션 완전성과 동작 정확성을 모두 측정한다. (1) 마이그레이션 감사(Migration Audit)는 마이그레이션이 실제로 발생했는지 검증한다. (2) 동작 테스트(Behavioural Tests)는 고정된 테스트 스위트를 사용하여 정확성을 측정한다. (3) 에이전트 기반 검증(Agentic Verification)은 6개의 독립적인 코딩 에이전트를 활용하여 숨겨진 동작 차이를 찾기 위한 표적 테스트를 생성한다. 8개의 최첨단 모델과 26개의 모델-노력 구성에서 수행된 520회의 실행 중 단 28회(5.4%)만이 3단계를 모두 통과했으며, 20개 작업 중 13개는 허용 가능한 솔루션을 얻지 못했고, 최고 성능 모델(claude-opus-5)은 47.0/100점을 기록했다. 마이그레이션 완전성과 동작 정확성은 서로 다른 능력이다. 소수의 실행은 마이그레이션을 건너뜀으로써 동작을 보존하며 마이그레이션 감사 단계에서 중단되고, 대부분의 실행은 마이그레이션을 시도하지만 동작을 깨뜨려 동작 테스트 단계에서 중단된다. 에이전트는 완벽한 마이그레이션을 제공하지 못한다. 마이그레이션 감사를 통과한 340회의 실행 중 58%가 고정 검사의 99%에 도달하지만, 100%에 도달하는 실행은 26%에 불과하다. 에이전트의 능력은 마이그레이션 범주에 따라 상이하다. 에이전트는 빌드 툴체인 재작성에서 31.4점을 기록한 반면, 언어 재작성에서는 5.6점에 그쳤다. 이러한 결과는 SWE Refactor Bench를 신뢰할 수 있는 전체 저장소 마이그레이션을 위한 코딩 에이전트 개발의 엄밀한 테스트베드로 자리매김하게 한다.
English
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.