ASI-Bench:人工超級智能的黎明
ASI-Bench: At the Dawn of Artificial Superintelligence
August 18, 2026
作者: Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie
cs.AI
摘要
人工超級智慧(ASI)要求人工智慧超越對既有知識的掌握,轉向探索未知、創造新知識,並將新想法轉化為可驗證的成果。然而,當前人工智慧系統的能力仍主要建立在學習、壓縮與應用既有的人類知識之上。因此,現有基準主要測試人工智慧能否根據所學知識給出正確答案,或能否在大量人類指導下完成任務。為此,我們推出 ASI-Bench,這是首個在一般研究領域中同時評估人工智慧系統創新探索與自主科學執行能力的基準,也是首個在同一研究專案內逐步移除人類方法學指導,以測試人工智慧能獨立進展到何種程度的基準。ASI-Bench 由超過 40 位專家投入 31,000 以上工時建置,包含橫跨 11 個科學領域的 60 項專案級研究任務,並逐步減少方法學指導,以測試人工智慧能否獨立選擇方法、進行研究並產出可驗證的結果。所有任務皆經過專家審查、AI 輔助稽核、沙盒執行與評分者驗證。在 18 種最先進的代理-模型配置中,平均分數從完整方法學指導下的 50.91 分,降至僅指定方法時的 29.10 分,再降至代理須自行決定方法時的 26.62 分。這種急遽下降顯示當前系統仍高度依賴人類指導,離自主執行端對端、專案級科學研究尚有極大差距。ASI-Bench 現已向全球開放。我們邀請各地的研究人員與開發者貢獻新任務、挑戰當前 AI 的極限,並協助加速人類邁向人工超級智慧的集體進程,請參閱 https://asibench.apexin.ai/submit。
English
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.