Hugging Face Daily Papers · · 3 min read

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Code: <a href=\"https://github.com/OpenMOSS/SWE-bench-Science\" rel=\"nofollow\">https://github.com/OpenMOSS/SWE-bench-Science</a><br>Data: <a href=\"https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science\">https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science</a><br>Leaderboard: <a href=\"https://swescience.github.io\" rel=\"nofollow\">https://swescience.github.io</a></p>\n","updatedAt":"2026-08-21T05:54:14.220Z","author":{"_id":"6729d2c902305fd2366b3763","avatarUrl":"/avatars/d957eee957d5d53d1915972dae15a6ef.svg","fullname":"yxzwang (SII)","name":"yxzwang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.641990065574646},"editors":["yxzwang"],"editorAvatarUrls":["/avatars/d957eee957d5d53d1915972dae15a6ef.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.19799","authors":[{"_id":"6a87e5d189e517cbfd75dc98","name":"Zhipeng Xu","hidden":false},{"_id":"6a87e5d189e517cbfd75dc99","name":"Jiahao Lu","hidden":false},{"_id":"6a87e5d189e517cbfd75dc9a","name":"Yining Zheng","hidden":false},{"_id":"6a87e5d189e517cbfd75dc9b","name":"Yuxin Wang","hidden":false},{"_id":"6a87e5d189e517cbfd75dc9c","name":"Xipeng Qiu","hidden":false}],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?","submittedOnDailyBy":{"_id":"6729d2c902305fd2366b3763","avatarUrl":"/avatars/d957eee957d5d53d1915972dae15a6ef.svg","isPro":false,"fullname":"yxzwang (SII)","user":"yxzwang","type":"user","name":"yxzwang"},"summary":"Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce SWE-bench Science, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, Claude Code with Opus-5 (max), achieves a pass@1 below 50\\%, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.","upvotes":31,"discussionId":"6a87e5d289e517cbfd75dc9d","projectPage":"https://swescience.github.io","githubRepo":"https://github.com/OpenMOSS/SWE-bench-Science","githubRepoAddedBy":"user","ai_summary":"SWE-bench Science benchmarks coding agents on scientific software repair, revealing failure mechanisms and mixed effects of scientific guidance.","ai_keywords":["SWE-bench Science","scientific software engineering","coding agents","pass@1","failure mechanisms","scientific knowledge","ablation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":13,"organization":{"_id":"613b0dee83ec35d460684607","name":"OpenMOSS-Team","fullname":"OpenMOSS","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61457b8deff2c9fdb4de4988/N5b9663zQ4uq5_OTNlnmw.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"667cced9cb6800a191427c1f","avatarUrl":"/avatars/9802f6f6eefcc98de89fda29860e8000.svg","isPro":false,"fullname":"Zhen Yu","user":"ZaneYue","type":"user"},{"_id":"6783634023394fe0bccafe65","avatarUrl":"/avatars/ae96cacb2a334e7e5af06c1ad4f58ea6.svg","isPro":false,"fullname":"Bowen Li","user":"Semophore","type":"user"},{"_id":"6511a8616f99a9540244c20f","avatarUrl":"/avatars/c3af9bacb050eea0e6a38d515e86a821.svg","isPro":false,"fullname":"Changsong","user":"maybe1possible","type":"user"},{"_id":"6729d2c902305fd2366b3763","avatarUrl":"/avatars/d957eee957d5d53d1915972dae15a6ef.svg","isPro":false,"fullname":"yxzwang (SII)","user":"yxzwang","type":"user"},{"_id":"67d27956919e3b981ac85295","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67d27956919e3b981ac85295/7aH8r7lHuFGnGafFMfyF_.jpeg","isPro":false,"fullname":"Wenxuan Wang","user":"wx9Songs","type":"user"},{"_id":"6539245ef940c8a03593f0fd","avatarUrl":"/avatars/617467ec59c9b4e76f39474e03cc9e51.svg","isPro":false,"fullname":"Xiaomeng Qian","user":"Qxm-bot","type":"user"},{"_id":"690ac605e7497f3dbbf8454e","avatarUrl":"/avatars/6ade69bb33e80f1aade796e7b1bb888b.svg","isPro":false,"fullname":"WendyCheung","user":"WendyCheung","type":"user"},{"_id":"64b74d1000bac1088ce92bee","avatarUrl":"/avatars/0ea6f2d65a8add7eeb3b3c1d84f16a05.svg","isPro":false,"fullname":"CHUAN YUAN TAN","user":"Cytan17726","type":"user"},{"_id":"68bbf3e203a1179f02eeccf2","avatarUrl":"/avatars/1e94e2ada1e0718e3987dfc3b6c8316a.svg","isPro":false,"fullname":"fang(SII)","user":"ign1s","type":"user"},{"_id":"680f7d6b8b2e2c7db910962c","avatarUrl":"/avatars/99d78c3c7d04a79121409c10d84df83a.svg","isPro":false,"fullname":"huazzeng","user":"huazzeng","type":"user"},{"_id":"6728aeb93a5bc303df90265f","avatarUrl":"/avatars/7f4fb0eb5ce13c86a3caf7ff9b79ff0d.svg","isPro":false,"fullname":"陈嘉骏","user":"ljbro","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"613b0dee83ec35d460684607","name":"OpenMOSS-Team","fullname":"OpenMOSS","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61457b8deff2c9fdb4de4988/N5b9663zQ4uq5_OTNlnmw.png"},"query":{}}">
Papers
arxiv:2608.19799

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Published on Aug 20
· Submitted by
yxzwang (SII)
on Aug 21
Authors:
,

Abstract

SWE-bench Science benchmarks coding agents on scientific software repair, revealing failure mechanisms and mixed effects of scientific guidance.

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce SWE-bench Science, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, Claude Code with Opus-5 (max), achieves a pass@1 below 50\%, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.19799 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.19799 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.19799 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers