Everyone says AI agents can already analyze data, run code, generate figures, and write reports.</p>\n<p>But can they actually complete a scientific workflow?</p>\n<p>Today, we’re introducing FrontierChallenge, a new benchmark evaluating whether AI agents can complete real scientific workflows end to end and deliver complete, verifiable results.</p>\n<p>We evaluated 12 frontier models across 97 cross-domain tasks.</p>\n<p>The highest full-completion rate was <strong>only 20.6% (GPT-5.6 Sol + Codex and Grok 4.6 + Claude Code).</strong></p>\n<p>In electrochemistry and environmental science, every evaluated system achieved a <strong>0% pass rate.</strong></p>\n<p>More strikingly, 75.5% of unsuccessful Claude Code runs still ended by claiming completion.</p>\n<p>Saying “done” is not the same as delivering.</p>\n<p>Agents that can advance scientific work are already here. Agents that can reliably complete scientific workflows are NOT.</p>\n<p>That’s why we built FrontierChallenge.</p>\n<p>🏆 Leaderboard<br><a href=\"https://apodexai.github.io/FrontierAgent/benchmarks/FrontierChallenge/\" rel=\"nofollow\">https://apodexai.github.io/FrontierAgent/benchmarks/FrontierChallenge/</a></p>\n<p>💻 GitHub<br><a href=\"https://github.com/ApodexAI/FrontierAgent/tree/main/benchmarks/frontierchallenge\" rel=\"nofollow\">https://github.com/ApodexAI/FrontierAgent/tree/main/benchmarks/frontierchallenge</a></p>\n<p>🤗 Hugging Face<br><a href=\"https://huggingface.co/datasets/apodex/FrontierChallenge\">https://huggingface.co/datasets/apodex/FrontierChallenge</a></p>\n<p>📝 Blog<br><a href=\"https://www.apodex.com/blog/frontier-challenge-evaluating-scientific-workflow-completion\" rel=\"nofollow\">https://www.apodex.com/blog/frontier-challenge-evaluating-scientific-workflow-completion</a></p>\n","updatedAt":"2026-08-27T02:26:56.683Z","author":{"_id":"677f945ea82c316db164a180","avatarUrl":"/avatars/50ec99d971564944de3b1d9c17d50cfd.svg","fullname":"Liangcai Su","name":"HKU-Liangcai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7927458882331848},"editors":["HKU-Liangcai"],"editorAvatarUrls":["/avatars/50ec99d971564944de3b1d9c17d50cfd.svg"],"reactions":[{"reaction":"🔥","users":["HKU-Liangcai"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24979","authors":[{"_id":"6a8f9fc42c24e8c5fab3291d","user":{"_id":"677f945ea82c316db164a180","avatarUrl":"/avatars/50ec99d971564944de3b1d9c17d50cfd.svg","isPro":false,"fullname":"Liangcai Su","user":"HKU-Liangcai","type":"user","name":"HKU-Liangcai"},"name":"Liangcai Su","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.927Z","hidden":false},{"_id":"6a8f9fc42c24e8c5fab3291e","name":"Zhaopeng Feng","hidden":false},{"_id":"6a8f9fc42c24e8c5fab3291f","name":"Zhuo Chen","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32920","name":"Zhen Zhang","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32921","name":"Xiang Lin","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32922","name":"Ruilin Li","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32923","name":"Handuo Zhang","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32924","name":"Ning Wang","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32925","name":"Kailong Wen","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32926","name":"Yueqi Guo","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32927","name":"Feng Xing","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32928","name":"Yiling Guo","hidden":false},{"_id":"6a8f9fc42c24e8c5fab32929","name":"Chenxiong Qian","hidden":false},{"_id":"6a8f9fc42c24e8c5fab3292a","name":"Simon Shaolei Du","hidden":false},{"_id":"6a8f9fc42c24e8c5fab3292b","name":"Lidong Bing","hidden":false},{"_id":"6a8f9fc42c24e8c5fab3292c","name":"Xinyu Wang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/677f945ea82c316db164a180/VJ7yJWXc2mhdVKTjmLd_O.mp4"],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"FrontierChallenge: Evaluating Scientific Workflow Completion","submittedOnDailyBy":{"_id":"677f945ea82c316db164a180","avatarUrl":"/avatars/50ec99d971564944de3b1d9c17d50cfd.svg","isPro":false,"fullname":"Liangcai Su","user":"HKU-Liangcai","type":"user","name":"HKU-Liangcai"},"summary":"Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.","upvotes":97,"discussionId":"6a8f9fc42c24e8c5fab3292d","projectPage":"https://apodexai.github.io/FrontierAgent/benchmarks/FrontierChallenge/","ai_summary":"FrontierChallenge evaluates end-to-end scientific workflows across domains, revealing that frontier models complete only about 20% of tasks despite high partial scores and frequent claims of completion.","ai_keywords":["agent scaffolds","Pass Rate","Avg. Score","scientific deliverables","end-to-end workflow execution"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"69e06d3f92a1019939f7e7a0","name":"apodex","fullname":"Apodex","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68bf96f0ca298a370d66adbb/T4Bt5ELjozQxgkKsHM_gg.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"677f945ea82c316db164a180","avatarUrl":"/avatars/50ec99d971564944de3b1d9c17d50cfd.svg","isPro":false,"fullname":"Liangcai Su","user":"HKU-Liangcai","type":"user"},{"_id":"63c0bb2a3bdc86f8108c112a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c0bb2a3bdc86f8108c112a/QvZlSEnyud-UR4p1SaGCk.jpeg","isPro":false,"fullname":"Zhaopeng Feng","user":"fzp0424","type":"user"},{"_id":"628c6240e410cb05ea07db20","avatarUrl":"/avatars/d6f4321fffe50207ec9680ce6bce235e.svg","isPro":false,"fullname":"Ziven","user":"SugaryLeon","type":"user"},{"_id":"64b73f9317570fdff9b0d1c4","avatarUrl":"/avatars/62124ae3e929b53f99f37e97226a877d.svg","isPro":false,"fullname":"Wang Xinyu","user":"oriuta","type":"user"},{"_id":"644e3e5f030210812f413073","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/uW8TKV2sds97lFBwnt6JK.jpeg","isPro":true,"fullname":"Zilong Chen","user":"heheyas","type":"user"},{"_id":"6466e7be1343dce20e59191b","avatarUrl":"/avatars/82b3796946127347d04b4a7f8dfc315a.svg","isPro":false,"fullname":"Li Ruilin","user":"Eric-LRL-130","type":"user"},{"_id":"641129818573c51c0458b793","avatarUrl":"/avatars/d4bc67c160a07146cf41c614678aa36b.svg","isPro":false,"fullname":"Tianqing Fang","user":"tqfang229","type":"user"},{"_id":"69ba304a13d2153b11def97d","avatarUrl":"/avatars/e81d608ef3f7656473d652c25847cc72.svg","isPro":false,"fullname":"Rock","user":"new-rock","type":"user"},{"_id":"68bf96f0ca298a370d66adbb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/xKpsUJqktcY_KbhFHlj6g.png","isPro":false,"fullname":"JohnnyIse","user":"IseJoe","type":"user"},{"_id":"64192ab522270b3ccf175d38","avatarUrl":"/avatars/3fec0b0216c8e1e728550742090d4de5.svg","isPro":false,"fullname":"Qi Fu","user":"qifu1991","type":"user"},{"_id":"63f43db70be81bdc5d9b0270","avatarUrl":"/avatars/6f48cf376edb9006290e8b958f2ea802.svg","isPro":false,"fullname":"Chuhui","user":"chxue233","type":"user"},{"_id":"650488e454b989666d042a49","avatarUrl":"/avatars/3dc79c6f1a9dce872636dddd38a04670.svg","isPro":false,"fullname":"Jiacheng Lin","user":"linjc16","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"69e06d3f92a1019939f7e7a0","name":"apodex","fullname":"Apodex","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68bf96f0ca298a370d66adbb/T4Bt5ELjozQxgkKsHM_gg.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24979.md","query":{}}">
FrontierChallenge: Evaluating Scientific Workflow Completion
Abstract
FrontierChallenge evaluates end-to-end scientific workflows across domains, revealing that frontier models complete only about 20% of tasks despite high partial scores and frequent claims of completion.
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.24979 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.24979 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.