FrontierChallenge:科学的ワークフロー完了の評価

FrontierChallenge: Evaluating Scientific Workflow Completion

August 25, 2026
著者: Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang
cs.AI

要旨

科学エージェントは、データ解析、コード実行、研究成果物の生成をますます行うようになっているが、既存のベンチマークの大半は、最終回答、独立したプログラム、あるいは単一の領域に重点を置いている。我々は、300件のエンドツーエンドの科学ワークフローからなるクロスドメインベンチマークであるFrontierChallengeを導入する。本論文では、量子化学、分子動力学、材料特性評価、分析化学、ライフサイエンス、電気化学/環境にわたる97タスクを公開し、評価する。各タスクは固定された入力と、必要とされる科学的成果物一式を規定する。我々は、3つのエージェントスキャフォールドを用いて12のフロンティアモデルを評価する。Pass Rateは完全達成基準を満たすタスクの割合を測定し、Avg. Scoreは部分的な進捗を捉える。最良の設定のそれぞれでも、公開された97タスクのうち完了できたのは20タスクのみで、Pass Rateは20.6%にとどまった。部分的な進捗は、分析化学および電気化学/環境において特に完全な成果物納品への変換が不十分であり、Avg. Scoreは87.6と94.9に達したものの、最高のPass Rateはそれぞれ4%と0%にすぎなかった。Passに至らなかったClaude Codeの軌跡のうち75.5%は、それでも完了を主張する文言で終わっていた。これらの結果は、高い部分スコアも、完了を確信的に主張する文言も、科学的タスクが完全に納品されたことを確実には示さないことを明らかにしており、エンドツーエンドのワークフロー実行と科学的成果物の完全性を併せて評価する必要性を浮き彫りにしている。
English
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.
PDF1301August 28, 2026