想法具有基因組:科學譜系推理與譜系基礎想法生成之基準測試
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
July 9, 2026
作者: Yifan Zhou, Qihao Yang, Yan Li, Donggang Li, Xiru Hu, Hokin Deng, Ziyang Gong, Xuanyi Zhou, Huacan Wang, Xiangchao Yan, Wanghan Xu, Wenlong Zhang, Shaofeng Zhang, Yue Zhou, Yifan Yang, Zhihang Zhong, Xue Yang
cs.AI
摘要
科學概念很少從零開始。它們繼承機制、修補已知限制、重組先前工作的片段,與生物基因組極為相似。目前的基準測試仍很少能說明AI系統是否能遵循這種繼承結構。我們提出IdeaGene-Bench(IG-Bench),一個針對科學譜系推理與基於譜系的構想生成的基準。IG-Bench圍繞IdeaGene框架組織:每篇論文或提案被表示為一組最小化、帶類型、以證據為基礎的Idea Genome物件,而GenomeDiff對齊這些物件,以在六種操作演化動力學下記錄繼承、突變、遺失、外部導入和新穎插入。該基準包含1,961條黃金譜系軌跡、1,085個經策劃的Idea Genome物件,以及920個成對GenomeDiff記錄,涵蓋10個科學領域。它支援兩種評估。IG-Exam(42種任務類型,1,029個實例)測試閉式譜系推理,涵蓋Idea Genome抽象、繼承追蹤、演化推理和譜系驗證。IG-Arena則以基於譜系的群體演化分數(PES)評估生成能力,檢驗一個提案能否作為給定譜系群體的連貫後代插入:它應繼承正確的Idea Genome物件、與鄰近工作有意義地差異化,並為未來研究提供選擇價值。在14個基於LLM的科學家系統上的實驗揭示了一個組合性瓶頸。最強的系統在譜系推理上僅達到27.3%的精確準確率,且結構化的譜系背景會重新洗牌系統排名,而非均勻地幫助所有參與者。
English
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.