アイデアはゲノムを持つ:科学的系統推論と系統に基づくアイデア生成のベンチマーク
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
July 9, 2026
著者: Yifan Zhou, Qihao Yang, Yan Li, Donggang Li, Xiru Hu, Hokin Deng, Ziyang Gong, Xuanyi Zhou, Huacan Wang, Xiangchao Yan, Wanghan Xu, Wenlong Zhang, Shaofeng Zhang, Yue Zhou, Yifan Yang, Zhihang Zhong, Xue Yang
cs.AI
要旨
科学的なアイデアが白紙の状態から始まることはほとんどない。それらは、生物のゲノムと同様に、既存の仕組みを受け継ぎ、既知の限界を修復し、先行研究の断片を再結合する。しかし、現在のベンチマークでは、AIシステムがこの継承構造に従えるかどうかを評価することはほとんどできない。本稿では、科学的な系統推論および系統に基づくアイデア生成のためのベンチマークであるIdeaGene-Bench(IG-Bench)を提案する。IG-BenchはIdeaGeneフレームワークに基づいて構成されている。各論文や提案は、最小限で型付けされ、エビデンスに基づくIdea Genomeオブジェクトの集合として表現され、GenomeDiffがこれらのオブジェクトを整列させ、6つの操作的な進化ダイナミクス(継承、突然変異、喪失、外部からの導入、新規挿入)を記録する。本ベンチマークは、10の科学領域にわたって1,961件のゴールデン系統トレース、1,085個の厳選されたIdea Genomeオブジェクト、920件のペアワイズGenomeDiffレコードを含む。2種類の評価をサポートする。IG-Exam(42タスクタイプ、1,029インスタンス)は、Idea Genomeの抽象化、継承追跡、進化推論、系統検証にわたるクローズドフォームの系統推論をテストする。IG-Arenaは、系統条件付き集団進化スコア(PES)を用いて生成を評価し、提案が与えられた系統集団の一貫した子孫として挿入可能かどうかを問う。すなわち、適切なIdea Genomeオブジェクトを継承し、近接する研究から有意に変異し、将来の研究に対して選択的価値を提供する必要がある。14のLLMベースの科学者による実験では、構成上のボトルネックが明らかになった。最も強力なシステムでも、系統推論における完全一致精度は27.3%にとどまり、構造化された系統コンテキストは、すべての参加者に一様に利益をもたらすのではなく、システムのランキングを再編成した。
English
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.