The ultimate test for coding agents isn't local editing — it's whole-repo evolution, and right now, the survival rate is 5.4%.</p>\n<p>Today we’re releasing SWE Refactor Bench, a benchmark for long-horizon, whole-repository software stack migration.</p>\n<p>Coding agents are getting very good at fixing bugs.<br>But can they refactor an entire system, C → Rust, Maven → Gradle, POSIX → WebAssembly?</p>\n<p>We built 20 real migrations across projects, including SQLite, zlib, libsodium, and GraphHopper.</p>\n<p>520 runs. Only 28 survived all 3 stages. 13/20 tasks were solved by nobody.</p>\n<p>System-scale migration is still wide open.</p>\n<p>🏠 Homepage<br><a href=\"https://lab.einsia.ai/swe-refactor-bench/\" rel=\"nofollow\">https://lab.einsia.ai/swe-refactor-bench/</a></p>\n<p>🏆 Leaderboard<br><a href=\"https://lab.einsia.ai/swe-refactor-bench/leaderboard/\" rel=\"nofollow\">https://lab.einsia.ai/swe-refactor-bench/leaderboard/</a></p>\n<p>💻 GitHub<br><a href=\"https://github.com/Einsia/SWE-Refactor-Bench\" rel=\"nofollow\">https://github.com/Einsia/SWE-Refactor-Bench</a></p>\n","updatedAt":"2026-08-27T09:45:26.225Z","author":{"_id":"69e0e87514a8e29a1cf75605","avatarUrl":"/avatars/4de5fbd56fb2d6f45ca40fb33ebf9d85.svg","fullname":"Deyao Hong","name":"hongdy22","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8626746535301208},"editors":["hongdy22"],"editorAvatarUrls":["/avatars/4de5fbd56fb2d6f45ca40fb33ebf9d85.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23564","authors":[{"_id":"6a8fb4052c24e8c5fab329d7","user":{"_id":"69e0e87514a8e29a1cf75605","avatarUrl":"/avatars/4de5fbd56fb2d6f45ca40fb33ebf9d85.svg","isPro":false,"fullname":"Deyao Hong","user":"hongdy22","type":"user","name":"hongdy22"},"name":"Deyao Hong","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.974Z","hidden":false},{"_id":"6a8fb4052c24e8c5fab329d8","name":"Yizhe Chi","hidden":false},{"_id":"6a8fb4052c24e8c5fab329d9","name":"Wenyi Li","hidden":false},{"_id":"6a8fb4052c24e8c5fab329da","name":"Xiaoqiu Wang","hidden":false},{"_id":"6a8fb4052c24e8c5fab329db","name":"Mingju Gao","hidden":false},{"_id":"6a8fb4052c24e8c5fab329dc","name":"Kaisen Yang","hidden":false},{"_id":"6a8fb4052c24e8c5fab329dd","name":"Bingxiang He","hidden":false},{"_id":"6a8fb4052c24e8c5fab329de","name":"Youjie Zheng","hidden":false},{"_id":"6a8fb4052c24e8c5fab329df","name":"Calvin Xiao","hidden":false},{"_id":"6a8fb4052c24e8c5fab329e0","name":"Qinhuai Na","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/69e0e87514a8e29a1cf75605/fRtPt7hFqannPFAls4oGs.png"],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?","submittedOnDailyBy":{"_id":"69e0e87514a8e29a1cf75605","avatarUrl":"/avatars/4de5fbd56fb2d6f45ca40fb33ebf9d85.svg","isPro":false,"fullname":"Deyao Hong","user":"hongdy22","type":"user","name":"hongdy22"},"summary":"Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.","upvotes":9,"discussionId":"6a8fb4062c24e8c5fab329e1","projectPage":"https://lab.einsia.ai/swe-refactor-bench/","githubRepo":"https://github.com/Einsia/SWE-Refactor-Bench","githubRepoAddedBy":"user","ai_summary":"The study introduces a benchmark for evaluating autonomous software migration by coding agents, finding that current models rarely complete migrations correctly.","ai_keywords":["SWE Refactor Bench","migration audit","behavioural tests","agentic verification","coding agents","technical debt","whole-repository migrations"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":22},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69e0e87514a8e29a1cf75605","avatarUrl":"/avatars/4de5fbd56fb2d6f45ca40fb33ebf9d85.svg","isPro":false,"fullname":"Deyao Hong","user":"hongdy22","type":"user"},{"_id":"68298606b6f6b5c3fe8549e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68298606b6f6b5c3fe8549e2/YHmOIttwvUM4DtqFxDW9w.jpeg","isPro":false,"fullname":"Chunyu Liu","user":"akpon900","type":"user"},{"_id":"69e0ea5ef8ecb1586f7044c3","avatarUrl":"/avatars/a8fa0af0337b1ae4ac8ac332ca4e11c8.svg","isPro":false,"fullname":"Hanxi Xiao","user":"CrazyCalvin","type":"user"},{"_id":"68fdf676e39d2cbbb5a4316f","avatarUrl":"/avatars/6a8f7a2e0178e6aba2d3c9fa8181462c.svg","isPro":false,"fullname":"Qingle Liu","user":"Aoraku","type":"user"},{"_id":"63d4ac10640bb0f771805e50","avatarUrl":"/avatars/da2590965d6223c79698958db2b25315.svg","isPro":false,"fullname":"Bowen Wang","user":"abmfy","type":"user"},{"_id":"6a86d1cce25551f0e4aad11e","avatarUrl":"/avatars/4c8edce8d43d258b9f3abe88a5223a6c.svg","isPro":false,"fullname":"OliviaYi","user":"OliviaYii","type":"user"},{"_id":"6a9010be4933313e5562fc13","avatarUrl":"/avatars/2060f12bad095d45ba5a61bfb2d4adbc.svg","isPro":false,"fullname":"Alyx","user":"Azulzz","type":"user"},{"_id":"69a4635e23713679f51d8cf8","avatarUrl":"/avatars/18108764bc593e1f328e1b784334528d.svg","isPro":false,"fullname":"yong yan","user":"yany24","type":"user"},{"_id":"6a9018c22aa458bc538afac3","avatarUrl":"/avatars/57acdc64c06aff5877c0eb8e13b072e0.svg","isPro":false,"fullname":"selina","user":"selina5566","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23564.md","query":{}}">
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
Abstract
The study introduces a benchmark for evaluating autonomous software migration by coding agents, finding that current models rarely complete migrations correctly.
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
Community
The ultimate test for coding agents isn't local editing — it's whole-repo evolution, and right now, the survival rate is 5.4%.
Today we’re releasing SWE Refactor Bench, a benchmark for long-horizon, whole-repository software stack migration.
Coding agents are getting very good at fixing bugs.
But can they refactor an entire system, C → Rust, Maven → Gradle, POSIX → WebAssembly?
We built 20 real migrations across projects, including SQLite, zlib, libsodium, and GraphHopper.
520 runs. Only 28 survived all 3 stages. 13/20 tasks were solved by nobody.
System-scale migration is still wide open.
🏠 Homepage
https://lab.einsia.ai/swe-refactor-bench/
🏆 Leaderboard
https://lab.einsia.ai/swe-refactor-bench/leaderboard/
💻 GitHub
https://github.com/Einsia/SWE-Refactor-Bench
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.23564 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.23564 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.23564 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.