FrontierChallenge: Evaluating Scientific Workflow Completion

Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang

arXiv:2608.24979 · 2026-08-27 공개 · arXiv · PDF

task-completion frontier-models molecular-dynamics cross-domain-benchmark pass-rate scientific-workflow quantum-chemistry analytical-chemistry

Abstract

Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.

한국어 요약

한 줄 요약

FrontierChallenge는 과학적 워크플로우 완료 능력을 평가하는 다분야 벤치마크로, 97개 과제에서 최고 성능 모델의 Pass Rate가 20.6%에 불과한 것으로 나타났다.

핵심 기여도

핵심 아이디어

기존 과학적 평가 기준은 단일 답변, 단일 프로그램, 또는 특정 분야에 집중되어 있다. 이에 반해 **FrontierChallenge**는 과학 워크플로우의 **종단간 실행**과 **배포물 완전성**을 평가하는 새로운 기준을 제시한다. 이 연구는 단순히 결과를 도출하는 것이 아니라, **입력 처리부터 최종 배포물까지 일관된 과학적 결과물을 생성하는 능력**을 중시한다. 각 과제는 고정된 입력과 **필수 배포물 계약**(deliverable contract)을 포함하며, **Grader**를 통해 평가된다. 이는 과학적 신뢰성을 확보하기 위해 **검증 가능하고 재현 가능한 분석**이 필요하다는 점에서 핵심적 통찰이다.

기술적 접근법

주요 결과

의의 및 한계

FrontierChallenge는 과학적 워크플로우의 **종단간 실행**과 **배포물 일관성**을 평가하는 새로운 기준을 제시하며, 단순한 부분적 진행이나 완료 언급이 과학적 성공을 보장하지 않음을 보여준다. 이는 과학적 에이전트가 신뢰성을 확보하기 위해 **계약 준수**, **배포물 간 검증**, **증거 기반 완료 확인**이 필요하다는 점을 강조한다. 그러나 평가 범위는 **공개된 97개 과제**와 **특정 모델/구성**에 제한되며, **단일 실행** 기반으로 한 분석이라는 한계가 있다.

실용적 활용

FrontierChallenge는 **다분야 과학 연구 자동화**, **AI 기반 실험 설계**, **과학적 에이전트 성능 평가**에 활용 가능하다. 특히, **실험 재현성**, **결과 일관성**, **도구 활용 능력**을 평가하는 데 유용하며, **연구 개발 과정의 자동화** 및 **AI 도구 선택**에 기준을 제공한다.