ASI-Bench: 인공초지능의 여명에
ASI-Bench: At the Dawn of Artificial Superintelligence
August 18, 2026
저자: Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie
cs.AI
초록
인공 초지능(ASI)은 AI가 기존 지식을 습득하는 것을 넘어 미지의 영역을 탐험하고, 새로운 지식을 창출하며, 새로운 아이디어를 검증 가능한 결과로 전환할 것을 요구한다. 그러나 오늘날 AI 시스템의 역량은 여전히 주로 기존 인간 지식의 학습, 압축, 적용에 기반을 두고 있다. 이에 따라 기존 벤치마크들은 주로 AI가 학습된 지식을 바탕으로 정확한 답을 산출할 수 있는지, 또는 광범위한 인간의 안내 하에 과업을 완수할 수 있는지를 시험한다. 이에 우리는 일반 연구 분야 전반에 걸쳐 AI 시스템의 혁신적 탐험 능력과 자율적 과학 수행 능력을 공동으로 평가하는 최초의 벤치마크이자, 동일한 연구 프로젝트 내에서 인간의 방법론적 안내를 점진적으로 제거하여 AI가 어디까지 자율적으로 진행할 수 있는지를 시험하는 최초의 벤치마크인 ASI-Bench를 소개한다. 40명 이상의 전문가가 31,000시간 이상의 인력을 투입하여 구축한 ASI-Bench는 11개 과학 분야에 걸친 60개의 프로젝트 수준 연구 과업으로 구성되며, 방법론적 안내를 점진적으로 축소하여 AI가 방법을 독립적으로 선택하고, 연구를 수행하며, 검증 가능한 결과를 생산할 수 있는지를 시험한다. 모든 과업은 전문가 검토, AI 보조 감사, 샌드박스 실행, 스코어러 검증을 거친다. 18개의 최첨단 에이전트-모델 구성에 걸쳐, 평균 점수는 완전한 방법론적 안내가 있을 때 50.91에서, 방법만 지정되었을 때 29.10, 에이전트가 방법을 스스로 결정해야 할 때 26.62로 급락한다. 이러한 급격한 하락은 현재의 시스템이 여전히 인간의 안내에 크게 의존하고 있으며, 종단 간(end-to-end) 프로젝트 수준의 과학 연구를 자율적으로 수행하는 데는 아직 요원함을 보여준다. ASI-Bench는 전 세계에 공개되어 있다. 우리는 모든 곳의 연구자와 개발자들이 새로운 과업을 기여하고, 오늘날 AI의 한계에 도전하며, 인류의 인공 초지능을 향한 집단적 경로를 가속화하는 데 동참할 것을 초대한다: https://asibench.apexin.ai/submit.
English
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.