ChatPaper.aiChatPaper

ASI-Bench:人工超知能の黎明に

ASI-Bench: At the Dawn of Artificial Superintelligence

August 18, 2026
著者: Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie
cs.AI

要旨

人工超知能(ASI)には、AIが既存知識の習得を超え、未知を探求し、新たな知識を創出し、新しいアイデアを検証可能な成果へと変えていくことが求められる。しかし、今日のAIシステムの能力は依然として、既存の人間知識を学習・圧縮・適用することに大きく基づいている。したがって、既存のベンチマークは主に、AIが学習した知識に基づいて正しい回答を生成できるか、あるいは人間による広範なガイダンスの下でタスクを完了できるかを検証するものにとどまる。そこで我々は、ASI-Benchを導入する。これは、広範な研究領域におけるAIシステムの革新的探索能力と自律的科学的実行能力を同時に評価する初めてのベンチマークであり、同一研究プロジェクト内で人間の方法論的ガイダンスを段階的に撤廃し、AIがどこまで独力で進めるかを検証する初めてのベンチマークでもある。ASI-Benchは、40名以上の専門家によって31,000時間以上に及ぶ人的コストをかけて構築され、11の科学分野にわたる60のプロジェクトレベル研究タスクを含み、方法論的ガイダンスを段階的に削減することで、AIが方法を選択し、研究を実施し、検証可能な成果を生み出すことを独立して行えるかを検証する。すべてのタスクは、専門家によるレビュー、AI支援による監査、サンドボックス実行、スコアラーによる検証を経る。18の最先端エージェント・モデル構成全体で、平均スコアは、完全な方法論的ガイダンスが与えられた場合の50.91から、方法のみが指定された場合の29.10、エージェントが自ら方法を決定しなければならない場合の26.62へと低下する。この急激な低下は、現在のシステムが依然として人間のガイダンスに大きく依存しており、エンドツーエンドのプロジェクトレベル科学研究を自律的に実施するにはほど遠いことを示している。ASI-Benchは世界に公開されている。私たちは、世界中の研究者と開発者に対し、新たなタスクの提供、今日のAIの限界への挑戦、そして人工超知能へ向けた人類全体の歩みを加速させることへの協力を呼びかける。https://asibench.apexin.ai/submit
English
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.