FrontierChallenge:评估科学工作流完成度
FrontierChallenge: Evaluating Scientific Workflow Completion
August 25, 2026
作者: Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang
cs.AI
摘要
科学智能体日益能够分析数据、执行代码并产出研究工件,然而大多数基准测试仅关注最终答案、孤立程序或单一领域。我们提出FrontierChallenge,一个包含300个端到端科学工作流的跨领域基准。本文发布并评估了其中97项任务,涵盖量子化学、分子动力学、材料表征、分析化学、生命科学以及电化学/环境领域。每项任务提供固定输入,并规定一组必需的科学交付物。我们使用三种智能体脚手架评估了十二个前沿模型。通过率(Pass Rate)衡量满足完整完成标准的任务比例,而平均分(Avg. Score)则捕捉部分进展。每个性能最优的配置仅完成了97项已发布任务中的20项,通过率为20.6%。在分析化学和电化学/环境领域,部分进展尤其难以转化为完整交付:平均分分别达到87.6和94.9,但最高通过率仅为4%和0%。在未通过的Claude Code轨迹中,仍有75.5%以声称完成的语言结尾。这些发现表明,无论是高分部分得分还是自信的完成声明,都无法可靠地表明科学任务已被完整交付,凸显了对端到端工作流执行与科学交付物完整性进行联合评估的必要性。
English
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.