아이디어 유전체: 과학적 계통 추론 및 계통 기반 아이디어 생성의 벤치마킹
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
July 9, 2026
저자: Yifan Zhou, Qihao Yang, Yan Li, Donggang Li, Xiru Hu, Hokin Deng, Ziyang Gong, Xuanyi Zhou, Huacan Wang, Xiangchao Yan, Wanghan Xu, Wenlong Zhang, Shaofeng Zhang, Yue Zhou, Yifan Yang, Zhihang Zhong, Xue Yang
cs.AI
초록
과학적 아이디어는 백지 상태에서 시작되는 경우가 드물다. 그들은 생물학적 게놈과 마찬가지로 메커니즘을 물려받고, 알려진 한계를 수정하며, 이전 연구의 조각들을 재조합한다. 현재의 벤치마크는 AI 시스템이 이러한 상속 구조를 따를 수 있는지에 대해 거의 말해주지 않는다. 우리는 과학적 계통 추론(linage reasoning)과 계통 기반 아이디어 생성을 위한 벤치마크인 IdeaGene-Bench(IG-Bench)를 제시한다. IG-Bench는 IdeaGene 프레임워크를 중심으로 구성된다. 각 논문이나 제안은 최소한의 유형화된 증거 기반 아이디어 게놈 객체(Idea Genome objects)의 집합으로 표현되며, GenomeDiff는 이러한 객체를 정렬하여 여섯 가지 작동적 진화 역학(상속, 돌연변이, 손실, 외부 도입, 새로운 삽입) 하에서의 상속, 돌연변이, 손실, 외부 도입 및 새로운 삽입을 기록한다. 이 벤치마크는 10개의 과학 분야에 걸쳐 1,961개의 황금 계통 추적(golden lineage traces), 1,085개의 큐레이션된 아이디어 게놈 객체, 920개의 쌍별 GenomeDiff 기록을 포함한다. 이는 두 가지 평가를 지원한다. IG-Exam(42개 작업 유형, 1,029개 인스턴스)은 아이디어 게놈 추상화, 상속 추적, 진화 추론 및 계통 검증 전반에 걸린 폐쇄형 계통 추론을 테스트한다. IG-Arena는 계통 조건부 인구-진화 점수(Population-Evolution Score, PES)를 사용한 생성을 평가하며, 제안이 주어진 계통 모집단의 일관된 후손으로 삽입될 수 있는지 묻는다. 즉, 올바른 아이디어 게놈 객체를 상속하고, 주변 연구와 의미 있게 차별화되며, 미래 연구를 위한 선택 가치를 제공해야 한다. 14개의 LLM 기반 과학자에 대한 실험은 구성적 병목 현상을 드러낸다. 가장 강력한 시스템도 계통 추론에서 정확도가 27.3%에 불과하며, 구조화된 계통 맥락은 모든 참가자에게 균일하게 도움을 주기보다 시스템 순위를 재구성한다.
English
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.