Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang

arXiv:2608.11341 · 2026-08-17 공개 · arXiv · PDF

foundation-model real-world-benchmarks drug-repurposing apodex-discovery problem-scouting hds6-evaluation aav-capsid-design tool-verification

Abstract

Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.

한국어 요약

한 줄 요약

Apodex Discovery는 실제 문제 해결을 위한 AI 평가 프레임워크로, AAV 캡시드 설계에서 기존 최고 성능을 7% 초과 달성.

핵심 기여도

핵심 아이디어

Apodex Discovery는 단순히 정답을 맞히는 AI가 아닌, 새로운 사실을 발견하는 AI를 평가하기 위한 프레임워크이다. 기존 AI 평가 방식은 고정된 질문-정답 형식이나, 실제 문제는 명확한 성공 기준 없이 제시된다. 이에 따라 Apodex는 **heavy-duty solver**라는 시스템을 도입하여, 기초 모델, 하네스, 도구, 제어 정책을 결합해 **지속적이고 검증 가능한 조사**를 수행하도록 설계했다. 핵심 아이디어는 **TRACES**라는 실행 환경을 통해 문제를 구조화하고, **HDS6**라는 프로세스 평가 체계를 통해 과학적 탐구의 질을 측정하는 것이다. 예를 들어, AAV 캡시드 설계에서 GPT-5.6-sol은 **의료 환경**을 사용함으로써 평균 정규화 예측 점수가 7.6포인트 상승한 것으로 보인다.

기술적 접근법

주요 결과

의의 및 한계

Apodex Discovery는 AI 평가를 단순 정답 맞히기에서 **실제 발견을 위한 과학적 탐구 과정 평가**로 전환시킨다. 특히, **HDS6**는 최종 결과가 불확실하거나 지연된 상황에서도 평가가 가능하다는 점에서 혁신적이다. 그러나 현재까지의 모델 성능은 여전히 **0.588**에 불과하며, **의약, 생물학, AI 엔지니어링** 등 17개 실행 환경에서의 성능 향상 여력이 크다. 또한, **모델-하네스 호환성**이 평균 성능 차이에 큰 영향을 미친다는 점도 한계로 지적된다.

실용적 활용

Apodex Discovery는 **의약품 개발**, **바이오의학 연구**, **AI 엔지니어링** 등 고가치 문제 해결에 적용 가능하다. 특히, **TRACES**는 실제 실험 환경과 유사한 조건에서 AI의 탐구 능력을 평가할 수 있어, **신약 개발**, **바이러스 구조 예측**, **AI 모델 최적화** 등에서 활용 가능하다.