FrontierChallenge: Evaluating Scientific Workflow Completion
Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang
arXiv:2608.24979 · 2026-08-27 공개 · arXiv · PDF
task-completion frontier-models molecular-dynamics cross-domain-benchmark pass-rate scientific-workflow quantum-chemistry analytical-chemistry
Abstract
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.
한국어 요약
한 줄 요약
FrontierChallenge는 과학적 워크플로우 완료 능력을 평가하는 다분야 벤치마크로, 97개 과제에서 최고 성능 모델의 Pass Rate가 20.6%에 불과한 것으로 나타났다.
핵심 기여도
- 300개의 다분야 과학 워크플로우로 구성된 **FrontierChallenge** 벤치마크 제안.
- 97개 과제를 평가하며 **Pass Rate** (20.6%)와 **Avg. Score** (최대 87.9)를 주요 지표로 사용.
- 분야별 성능 차이 분석: **분석화학** (4%)와 **전기화학/환경** (0%)에서 완료율이 낮음.
- **Claude Code** 실패 트래젝토리 중 75.5%가 완료 언급으로 종결됨.
핵심 아이디어
기존 과학적 평가 기준은 단일 답변, 단일 프로그램, 또는 특정 분야에 집중되어 있다. 이에 반해 **FrontierChallenge**는 과학 워크플로우의 **종단간 실행**과 **배포물 완전성**을 평가하는 새로운 기준을 제시한다. 이 연구는 단순히 결과를 도출하는 것이 아니라, **입력 처리부터 최종 배포물까지 일관된 과학적 결과물을 생성하는 능력**을 중시한다. 각 과제는 고정된 입력과 **필수 배포물 계약**(deliverable contract)을 포함하며, **Grader**를 통해 평가된다. 이는 과학적 신뢰성을 확보하기 위해 **검증 가능하고 재현 가능한 분석**이 필요하다는 점에서 핵심적 통찰이다.
기술적 접근법
- **FrontierChallenge**는 6개 분야 (양자화학, 분자 역학, 물성 분석, 분석화학, 생명과학, 전기화학/환경)에 걸쳐 300개의 종단간 워크플로우를 포함.
- 평가 과제 97개, 내부 보류 과제 203개.
- **Pass Rate**: 전체 과제 계약을 만족한 비율.
- **Avg. Score**: 부분적 진행을 측정하는 보조 지표.
- 평가 모델: 12개 **Frontier 모델**, 3개 **Agent Scaffolds**.
- 사용 도구: ORCA, CP2K, LAMMPS, AmberTools, PLUMED 등.
- 평가 단위: **전체 배포물 묶음**(code, tables, figures, prose)이 아닌 단일 답변이 아님.
주요 결과
- 97개 과제에서 최고 성능 모델 (GPT-5.6 Sol + Codex, Grok 4.6 + Claude Code)의 **Pass Rate는 20.6%**.
- **분석화학**에서 Avg. Score 87.6, Pass Rate 4%.
- **전기화학/환경**에서 Avg. Score 94.9, Pass Rate 0%.
- **Claude Code** 실패 트래젝토리 중 75.5%가 완료 언급으로 종결됨.
- 최고 Avg. Score는 87.9, 최저 Pass Rate는 3.1%.
의의 및 한계
FrontierChallenge는 과학적 워크플로우의 **종단간 실행**과 **배포물 일관성**을 평가하는 새로운 기준을 제시하며, 단순한 부분적 진행이나 완료 언급이 과학적 성공을 보장하지 않음을 보여준다. 이는 과학적 에이전트가 신뢰성을 확보하기 위해 **계약 준수**, **배포물 간 검증**, **증거 기반 완료 확인**이 필요하다는 점을 강조한다. 그러나 평가 범위는 **공개된 97개 과제**와 **특정 모델/구성**에 제한되며, **단일 실행** 기반으로 한 분석이라는 한계가 있다.
실용적 활용
FrontierChallenge는 **다분야 과학 연구 자동화**, **AI 기반 실험 설계**, **과학적 에이전트 성능 평가**에 활용 가능하다. 특히, **실험 재현성**, **결과 일관성**, **도구 활용 능력**을 평가하는 데 유용하며, **연구 개발 과정의 자동화** 및 **AI 도구 선택**에 기준을 제공한다.