Hugging Face Daily Papers · · 6 min read

ASI-Bench: At the Dawn of Artificial Superintelligence

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

ASI-Bench evaluates AI agents' capabilities in innovative scientific exploration and autonomous project-level research across multiple domains.</p>\n","updatedAt":"2026-08-19T02:30:05.563Z","author":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":361,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8029451966285706},"editors":["taesiri"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg"],"reactions":[],"isReport":false}},{"id":"6a8525698b14a84b97dbc161","author":{"_id":"657157dc971de7383e01ebc9","avatarUrl":"/avatars/70a58d41bd4f86191205e916e4f6373e.svg","fullname":"Zhou Xueyang","name":"zhouxueyang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-08-19T03:39:21.000Z","type":"comment","data":{"edited":true,"hidden":false,"latest":{"raw":"**🚀 ASI-Bench: At the Dawn of Artificial Superintelligence**\n\n**ASI-Bench** — released by researchers from **Tsinghua, MIT, Harvard, CMU, the Flatiron Institute, Microsoft Research, and more** — is designed to measure one critical capability:\n\n🧠 **Scientific autonomy in AI systems.**\n\n🔬 **60 project-level research tasks** across **11 scientific fields**\n\n📉 **A first-of-its-kind B1 → B4 guidance gradient**\nHuman methodological guidance is progressively removed to test whether AI can **choose its own methods, conduct the research autonomously, and produce verifiable scientific results.**\n\n🤖 **18 frontier Agent × Model configurations evaluated**\nWhen methodological guidance is removed, average performance drops from **50.91 → 26.62**.\n\n🏆 Even the **best system reaches only 51.60** under autonomous research settings.\n\nThe message is clear:\n\n**Today’s AI is increasingly capable of solving known problems — but autonomous scientific discovery remains far from solved.**\n\n#ASIBench #ArtificialSuperintelligence #AIScientist #AIforScience #ScientificDiscovery #AutonomousAgents #Benchmark #Research\n","html":"<p><strong>🚀 ASI-Bench: At the Dawn of Artificial Superintelligence</strong></p>\n<p><strong>ASI-Bench</strong> — released by researchers from <strong>Tsinghua, MIT, Harvard, CMU, the Flatiron Institute, Microsoft Research, and more</strong> — is designed to measure one critical capability:</p>\n<p>🧠 <strong>Scientific autonomy in AI systems.</strong></p>\n<p>🔬 <strong>60 project-level research tasks</strong> across <strong>11 scientific fields</strong></p>\n<p>📉 <strong>A first-of-its-kind B1 → B4 guidance gradient</strong><br>Human methodological guidance is progressively removed to test whether AI can <strong>choose its own methods, conduct the research autonomously, and produce verifiable scientific results.</strong></p>\n<p>🤖 <strong>18 frontier Agent × Model configurations evaluated</strong><br>When methodological guidance is removed, average performance drops from <strong>50.91 → 26.62</strong>.</p>\n<p>🏆 Even the <strong>best system reaches only 51.60</strong> under autonomous research settings.</p>\n<p>The message is clear:</p>\n<p><strong>Today’s AI is increasingly capable of solving known problems — but autonomous scientific discovery remains far from solved.</strong></p>\n<p>#ASIBench #ArtificialSuperintelligence #AIScientist #AIforScience #ScientificDiscovery #AutonomousAgents #Benchmark #Research</p>\n","updatedAt":"2026-08-19T03:41:56.980Z","author":{"_id":"657157dc971de7383e01ebc9","avatarUrl":"/avatars/70a58d41bd4f86191205e916e4f6373e.svg","fullname":"Zhou Xueyang","name":"zhouxueyang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":3,"identifiedLanguage":{"language":"en","probability":0.7431628108024597},"editors":["zhouxueyang"],"editorAvatarUrls":["/avatars/70a58d41bd4f86191205e916e4f6373e.svg"],"reactions":[{"reaction":"🚀","users":["taesiri","sijiachen1","zjw49246","jiarx","zx10086","Apex-sun","yongchao98"],"count":7},{"reaction":"🔥","users":["taesiri","Apex-sun","yongchao98"],"count":3}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17271","authors":[{"_id":"6a85150f536bdd3bdd48f7a3","name":"Junwei Zhou","hidden":false},{"_id":"6a85150f536bdd3bdd48f7a4","name":"Zhen Sun","hidden":false},{"_id":"6a85150f536bdd3bdd48f7a5","name":"Binyu Li","hidden":false},{"_id":"6a85150f536bdd3bdd48f7a6","name":"Jiangyu Zhou","hidden":false},{"_id":"6a85150f536bdd3bdd48f7a7","name":"Yuexi Pan","hidden":false},{"_id":"6a85150f536bdd3bdd48f7a8","name":"Hengyu Wang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7a9","name":"Honghe Ren","hidden":false},{"_id":"6a85150f536bdd3bdd48f7aa","name":"Xiaohan Jia","hidden":false},{"_id":"6a85150f536bdd3bdd48f7ab","name":"Xueyang Zhou","hidden":false},{"_id":"6a85150f536bdd3bdd48f7ac","name":"Xiaoyu Cao","hidden":false},{"_id":"6a85150f536bdd3bdd48f7ad","name":"Yongchao Chen","hidden":false},{"_id":"6a85150f536bdd3bdd48f7ae","name":"Yuanning Feng","hidden":false},{"_id":"6a85150f536bdd3bdd48f7af","name":"Junhao Wu","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b0","name":"Cheng Zhang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b1","name":"Sijia Chen","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b2","name":"Haoyu Xue","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b3","name":"Chengsong You","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b4","name":"Huan Wang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b5","name":"Koutian Wu","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b6","name":"Peigan Gao","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b7","name":"Jiakun Wu","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b8","name":"Wenzhe Li","hidden":false},{"_id":"6a85150f536bdd3bdd48f7b9","name":"Ergan Shang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7ba","name":"Qingyuan Zheng","hidden":false},{"_id":"6a85150f536bdd3bdd48f7bb","name":"Jingjing Zhou","hidden":false},{"_id":"6a85150f536bdd3bdd48f7bc","name":"Ruixuan Jia","hidden":false},{"_id":"6a85150f536bdd3bdd48f7bd","name":"Yan Xu","hidden":false},{"_id":"6a85150f536bdd3bdd48f7be","name":"Hongrui Zhang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7bf","name":"Xiao-Han Ma","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c0","name":"Zhengxiang Cheng","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c1","name":"Yuexing Hao","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c2","name":"Liting Mai","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c3","name":"Xianglin Ji","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c4","name":"Wenjun Zhang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c5","name":"Zhuofan Chen","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c6","name":"Yixiao Huang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c7","name":"Chi Wang","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c8","name":"Wenyue Hua","hidden":false},{"_id":"6a85150f536bdd3bdd48f7c9","name":"Yilun Hao","hidden":false},{"_id":"6a85150f536bdd3bdd48f7ca","name":"Yuantao Zhai","hidden":false},{"_id":"6a85150f536bdd3bdd48f7cb","name":"Ziyan Zhao","hidden":false},{"_id":"6a85150f536bdd3bdd48f7cc","name":"Jingyan Xie","hidden":false}],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"ASI-Bench: At the Dawn of Artificial Superintelligence","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.","upvotes":38,"discussionId":"6a851510536bdd3bdd48f7cd","projectPage":"https://asibench.apexin.ai/","githubRepo":"https://github.com/apexin-ai/ASI-Bench","githubRepoAddedBy":"user","githubStars":10},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"657157dc971de7383e01ebc9","avatarUrl":"/avatars/70a58d41bd4f86191205e916e4f6373e.svg","isPro":false,"fullname":"Zhou Xueyang","user":"zhouxueyang","type":"user"},{"_id":"6662fb7009d721eaab7dde08","avatarUrl":"/avatars/e87c688a2a0c9db1f4667fe614ad037f.svg","isPro":false,"fullname":"Yuanning Feng","user":"plafle","type":"user"},{"_id":"68e9f334d436b990830b85af","avatarUrl":"/avatars/0ae14f07f5101838148fec1808f451d3.svg","isPro":false,"fullname":"Zhuofan Chen","user":"Vicky0719","type":"user"},{"_id":"669096da35cddb688a352ca8","avatarUrl":"/avatars/5dd096cb7360682016d0fca909ab9744.svg","isPro":false,"fullname":"zxiang","user":"zx10086","type":"user"},{"_id":"6a4daefc5ff747e27059370c","avatarUrl":"/avatars/9af0ce795ab3197935e6c3b2a6851dc5.svg","isPro":false,"fullname":"Jiangyu Zhou","user":"jvzhou","type":"user"},{"_id":"680051edc771c307fec5e889","avatarUrl":"/avatars/8d0396fa3cf002a02cd2984046e6f41c.svg","isPro":false,"fullname":"You","user":"Justin7219","type":"user"},{"_id":"6a85286f90a182a811ccd15c","avatarUrl":"/avatars/c55ff95c03d47c56f7eae8fa98dc1280.svg","isPro":false,"fullname":"caoxiaoyu","user":"caoxiaoyuyuyuyuyuyuyu","type":"user"},{"_id":"669ca072990749deca6e30db","avatarUrl":"/avatars/f4647f4d1f0df2a569551a399c1f63b5.svg","isPro":false,"fullname":"Ruixuan Jia","user":"jiarx","type":"user"},{"_id":"6902abbceadaabdf99c51a7d","avatarUrl":"/avatars/bef68724389db2e46d91210e6f39fe02.svg","isPro":false,"fullname":"Jingyan Xie","user":"Jean1120","type":"user"},{"_id":"6a798cc8b9eeaff8deae9eb8","avatarUrl":"/avatars/e0b4ffb4869c274725dfb7d06167b191.svg","isPro":false,"fullname":"kelvin xing","user":"kelvin01X","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17271.md","query":{}}">
Papers
arxiv:2608.17271

ASI-Bench: At the Dawn of Artificial Superintelligence

Published on Aug 18
· Submitted by
taesiri
on Aug 19
#1 Paper of the day
Authors:
,

Abstract

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.

Community

Paper submitter about 6 hours ago

ASI-Bench evaluates AI agents' capabilities in innovative scientific exploration and autonomous project-level research across multiple domains.

🚀 ASI-Bench: At the Dawn of Artificial Superintelligence

ASI-Bench — released by researchers from Tsinghua, MIT, Harvard, CMU, the Flatiron Institute, Microsoft Research, and more — is designed to measure one critical capability:

🧠 Scientific autonomy in AI systems.

🔬 60 project-level research tasks across 11 scientific fields

📉 A first-of-its-kind B1 → B4 guidance gradient
Human methodological guidance is progressively removed to test whether AI can choose its own methods, conduct the research autonomously, and produce verifiable scientific results.

🤖 18 frontier Agent × Model configurations evaluated
When methodological guidance is removed, average performance drops from 50.91 → 26.62.

🏆 Even the best system reaches only 51.60 under autonomous research settings.

The message is clear:

Today’s AI is increasingly capable of solving known problems — but autonomous scientific discovery remains far from solved.

#ASIBench #ArtificialSuperintelligence #AIScientist #AIforScience #ScientificDiscovery #AutonomousAgents #Benchmark #Research

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.17271
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.17271 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.17271 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers