FrontierChallenge:評估科學工作流完成
FrontierChallenge: Evaluating Scientific Workflow Completion
August 25, 2026
作者: Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang
cs.AI
摘要
科學智能體日益廣泛地分析數據、執行程式碼,並產出研究產物,然而多數基準測試僅著重於最終答案、孤立程式或單一領域。我們提出了 FrontierChallenge,一個跨領域基準測試,包含 300 個端對端科學工作流程。在本論文中,我們公開並評估了其中 97 項任務,涵蓋量子化學、分子動力學、材料表徵、分析化學、生命科學,以及電化學/環境等領域。每項任務皆提供固定輸入,並指定一組必需的科學交付成果。我們以三種智能體框架評估了十二個前沿模型。通過率(Pass Rate)衡量滿足完整完成標準的任務比例,而平均分數(Avg. Score)則反映部分進展。每個表現最佳的配置僅完成了 97 項已公開任務中的 20 項,通過率為 20.6%。部分進展轉化為完整交付的成效,在分析化學及電化學/環境領域尤為不佳:平均分數分別達到 87.6 和 94.9,但最高通過率僅為 4% 和 0%。在未通過的 Claude Code 軌跡中,有 75.5% 仍以聲稱完成的語言作結。這些發現表明,無論是高分部分進展,還是自信的完成聲明,都無法可靠地指示科學任務已完整交付,這凸顯了同時評估端對端工作流程執行與科學交付成果完整性的必要。
English
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.