NatureBench: Agentes de Codificação Podem Igualar o SOTA Publicado dos Artigos da Família Nature?
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
June 23, 2026
Autores: Yuru Wang, Lejun Cheng, Yuxin Zuo, Sihang Zeng, Bingxiang He, Che Jiang, Junlin Yang, Yuchong Wang, Kaikai Zhao, Weifeng Huang, Kai Tian, Zhenzhao Yuan, Jincheng Zhong, Weizhi Wang, Ning Ding, Bowen Zhou, Kaiyan Zhang
cs.AI
Resumo
Apresentamos o NatureBench, um benchmark interdisciplinar composto por 90 tarefas destiladas de publicações revisadas por pares da família Nature, projetado para avaliar se agentes de codificação de IA podem avançar além da reprodução em direção à descoberta em problemas científicos reais. O NatureBench é construído sobre o NatureGym, um pipeline automatizado que cria um ambiente containerizado padronizado e específico para cada tarefa a partir de um artigo fonte, abordando o problema de fragmentação de ambientes que limitou a credibilidade de benchmarks anteriores focados em agentes para pesquisa. Ao avaliar dez configurações de agentes de fronteira sob um protocolo rigoroso com busca na web desabilitada, constatamos que o modelo mais forte supera o estado da arte em apenas 17,8% das tarefas, considerando o critério g > 0,1. A análise dos caminhos metodológicos revela que os agentes obtêm sucesso principalmente por meio de tradução metodológica, convertendo tarefas científicas em problemas familiares de predição supervisionada, em vez de invenção científica genuína. As falhas são dominadas por escolha inadequada de método e orçamento computacional insuficiente, e não por mal-entendido da tarefa. Disponibilizamos o benchmark, o pipeline NatureGym e um leaderboard público com reprodução pelo lado do mantenedor. Código: https://github.com/FrontisAI/NatureBench
English
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction toward discovery on real scientific problems. NatureBench is built on NatureGym, an automated pipeline that constructs a standardized, per-task containerized environment from a source paper, addressing the environment-fragmentation problem that has limited the credibility of prior agent-on-research benchmarks. Evaluating ten frontier agent configurations under a strict web-search-disabled protocol, we find that the strongest model surpasses SOTA on only 17.8% of tasks under the g>0.1 criterion. Analysis of method pathways reveals that agents succeed primarily through methodological translation, converting scientific tasks into familiar supervised prediction problems, rather than through genuine scientific invention. Failures are dominated by wrong method choice and insufficient compute budget, not by task misunderstanding. We release the benchmark, the NatureGym pipeline, and a public leaderboard with maintainer-side reproduction. Code: https://github.com/FrontisAI/NatureBench