FrontierChallenge: 과학적 워크플로우 완료 평가
FrontierChallenge: Evaluating Scientific Workflow Completion
August 25, 2026
저자: Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang
cs.AI
초록
과학적 에이전트는 점점 더 데이터를 분석하고 코드를 실행하며 연구 산출물을 생성하고 있지만, 대부분의 벤치마크는 최종 답변, 단독 프로그램, 또는 단일 도메인에 초점을 맞추고 있다. 우리는 300개의 종단 간(end-to-end) 과학 워크플로우로 구성된 교차 도메인 벤치마크인 FrontierChallenge를 소개한다. 본 논문에서는 양자 화학, 분자 동역학, 재료 특성 분석, 분석 화학, 생명과학, 전기화학/환경에 걸친 97개 과제를 공개하고 평가한다. 각 과제는 고정된 입력을 제공하고 요구되는 과학적 산출물 번들을 명시한다. 우리는 세 가지 에이전트 스캐폴드로 12개의 프런티어 모델을 평가한다. 통과율(Pass Rate)은 전체 완료 기준을 충족하는 과제의 비율을 측정하고, 평균 점수(Avg. Score)는 부분적 진전을 반영한다. 최고 성능을 보인 각 구성은 공개된 97개 과제 중 20개만 완료하여 20.6%의 통과율을 기록했다. 부분적 진전은 특히 분석 화학과 전기화학/환경 분야에서 완전한 전달로 이어지는 비율이 낮았다. 평균 점수는 각각 87.6과 94.9에 도달했지만, 최고 통과율은 각각 4%와 0%에 불과했다. 통과하지 못한 Claude Code 궤적 중 75.5%는 여전히 완료를 주장하는 문구로 종료되었다. 이러한 결과는 높은 부분 점수나 완료에 대한 확신에 찬 주장 모두 과학적 과제가 완전히 수행되었음을 신뢰성 있게 나타내지 않는다는 것을 보여주며, 종단 간 워크플로우 실행과 과학적 산출물의 완전성을 함께 평가할 필요성을 강조한다.
English
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.