ASI-Bench: At the Dawn of Artificial Superintelligence
Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie
arXiv:2608.17271 · 2026-08-19 공개 · arXiv · PDF
scientific-research ai-audit agent-model asi-bench artificial-superintelligence autonomous-execution method-guidance research-tasks
Abstract
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.
한국어 요약
한 줄 요약
ASI-Bench는 AI가 인간 지침 없이 독자적으로 과학 연구를 수행할 수 있는 능력을 평가하는 첫 번째 벤치마크로, 18개 최신 에이전트-모델 설정에서 평균 점수가 50.91에서 26.62로 급락함을 보여준다.
핵심 기여도
- ASI-Bench는 AI의 창의적 탐색과 자율적 과학 수행 능력을 종합적으로 평가하는 첫 번째 벤치마크이다.
- 60개 프로젝트 수준의 과학 연구 과제를 11개 분야에 걸쳐 구성하고, B1부터 B4까지 점차적으로 인간 지침을 줄이며 평가한다.
- 18개 최신 에이전트-모델 설정에서 평균 점수가 50.91 (B1) → 29.10 (B2) → 26.62 (B3)로 급락하며, AI의 자율성 부족을 입증한다.
- 모든 과제는 전문가 검토, AI 보조 감사, 샌드박스 실행, 점수 검증을 통해 과학적 신뢰성을 확보한다.
핵심 아이디어
기존 AI 시스템은 대부분 인간 지식을 학습하고 적용하는 데 기반을 두고 있으며, ASI(인공 초지능)는 AI가 새로운 지식을 창출하고 이를 검증 가능한 결과로 전환하는 능력이 필요하다. ASI-Bench는 AI가 인간 지침 없이 독자적으로 과학 연구를 수행할 수 있는지를 평가하기 위해 설계되었다. 이는 기존 벤치마크가 주로 알려진 문제나 명확한 절차를 기반으로 평가하는 한계를 극복하기 위한 시도이다.
ASI-Bench는 B1 (전체 지침), B2 (방법만 지정), B3 (방법 자율 결정), B4 (방법 자율 결정 + 방해 요소 추가)의 4단계 구조를 통해 AI의 자율성을 점진적으로 평가한다. 이 구조는 일반 지능, 창의성, 자율 실행 능력을 통합적으로 측정할 수 있도록 설계되었다.
기술적 접근법
- **데이터셋**: 11개 과학 분야에 걸쳐 60개의 프로젝트 수준 연구 과제 포함.
- **평가 구조**: B1 (전체 지침), B2 (방법만 지정), B3 (방법 자율 결정), B4 (방법 자율 결정 + 방해 요소 추가).
- **모델**: 18개 최신 에이전트-모델 설정으로 평가.
- **검증 과정**: 전문가 검토, AI 보조 감사, 샌드박스 실행, 점수 검증.
- **평가 지표**: 각 설정에서 평균 점수를 기준으로 성능을 비교.
- **하이퍼파라미터**: 명시되지 않음.
주요 결과
- 18개 최신 에이전트-모델 설정에서 평균 점수는 B1 (전체 지침)에서 50.91, B2 (방법만 지정)에서 29.10, B3 (방법 자율 결정)에서 26.62로 급락함.
- B3에서 B2 대비 -7.48점, B1 대비 -24.29점 감소.
- B4에서는 B3 대비 추가적인 방해 요소가 성능 저하를 유발함.
- AI는 인간 지침 없이 독자적으로 과학 연구를 수행하는 데 어려움을 보임.
의의 및 한계
ASI-Bench는 AI가 인간 지침 없이 과학 연구를 수행할 수 있는 능력을 평가하는 첫 번째 종합적 벤치마크로, AI의 자율성과 창의성을 측정하는 새로운 기준을 제시한다. 또한, AI가 ASI로 진화할 수 있는 여정을 추적하는 데 기초가 될 수 있다. 그러나 ASI-Bench는 고정된 벤치마크이므로, 미래 AI가 직면할 다양한 문제를 완전히 반영하기는 어렵다. 또한, 모든 과제가 샌드박스 환경에서 실행되므로 실제 연구 환경과의 차이가 있을 수 있다.
실용적 활용
ASI-Bench는 과학 연구 분야에서 AI의 자율성과 창의성을 평가하는 데 활용될 수 있으며, AI 연구자들이 새로운 모델과 에이전트를 개발하고 테스트하는 데 중요한 도구가 될 수 있다. 또한, AI 기반 연구 자동화 시스템의 개발 및 평가에도 활용 가능하다. 연구자와 엔지니어들이 공동으로 과제를 제출하고 벤치마크를 발전시킴으로써 AI의 ASI 진화를 가속화할 수 있다.