ASI-Bench: At the Dawn of Artificial Superintelligence

Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie

arXiv:2608.17271 · 2026-08-19 공개 · arXiv · PDF

scientific-research ai-audit agent-model asi-bench artificial-superintelligence autonomous-execution method-guidance research-tasks

Abstract

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.

한국어 요약

한 줄 요약

ASI-Bench는 AI가 인간 지침 없이 독자적으로 과학 연구를 수행할 수 있는 능력을 평가하는 첫 번째 벤치마크로, 18개 최신 에이전트-모델 설정에서 평균 점수가 50.91에서 26.62로 급락함을 보여준다.

핵심 기여도

핵심 아이디어

기존 AI 시스템은 대부분 인간 지식을 학습하고 적용하는 데 기반을 두고 있으며, ASI(인공 초지능)는 AI가 새로운 지식을 창출하고 이를 검증 가능한 결과로 전환하는 능력이 필요하다. ASI-Bench는 AI가 인간 지침 없이 독자적으로 과학 연구를 수행할 수 있는지를 평가하기 위해 설계되었다. 이는 기존 벤치마크가 주로 알려진 문제나 명확한 절차를 기반으로 평가하는 한계를 극복하기 위한 시도이다.

ASI-Bench는 B1 (전체 지침), B2 (방법만 지정), B3 (방법 자율 결정), B4 (방법 자율 결정 + 방해 요소 추가)의 4단계 구조를 통해 AI의 자율성을 점진적으로 평가한다. 이 구조는 일반 지능, 창의성, 자율 실행 능력을 통합적으로 측정할 수 있도록 설계되었다.

기술적 접근법

주요 결과

의의 및 한계

ASI-Bench는 AI가 인간 지침 없이 과학 연구를 수행할 수 있는 능력을 평가하는 첫 번째 종합적 벤치마크로, AI의 자율성과 창의성을 측정하는 새로운 기준을 제시한다. 또한, AI가 ASI로 진화할 수 있는 여정을 추적하는 데 기초가 될 수 있다. 그러나 ASI-Bench는 고정된 벤치마크이므로, 미래 AI가 직면할 다양한 문제를 완전히 반영하기는 어렵다. 또한, 모든 과제가 샌드박스 환경에서 실행되므로 실제 연구 환경과의 차이가 있을 수 있다.

실용적 활용

ASI-Bench는 과학 연구 분야에서 AI의 자율성과 창의성을 평가하는 데 활용될 수 있으며, AI 연구자들이 새로운 모델과 에이전트를 개발하고 테스트하는 데 중요한 도구가 될 수 있다. 또한, AI 기반 연구 자동화 시스템의 개발 및 평가에도 활용 가능하다. 연구자와 엔지니어들이 공동으로 과제를 제출하고 벤치마크를 발전시킴으로써 AI의 ASI 진화를 가속화할 수 있다.