大規模発見モデル:実証に基づくモデルベースのオープンエンド探索
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
August 16, 2026
著者: Zhongwei Yu, Yan Song, Xue Yan, Anjie Liu, Xingyu Lu, Yihang Chen, Huichi Zhou, Siyuan Guo, Luoyang Sun, Sihan Chen, Xiangning Yu, Jun Wang
cs.AI
要旨
科学的発見は、分子、タンパク質配列、コンピュータプログラムなど、広大で構造化され、開かれた仮説空間における評価コストの高い目的関数の最適化を伴うことが多い。大規模言語モデル(LLM)などの生成モデルは、そのような空間に対する表現力豊かな事前分布を提供するが、その尤度や自己評価は、目的関数や較正された認識論的不確実性の信頼できる代理とはなり得ない。特に、観測データ分布の外にある新規候補についてはなおさらである。本稿では、生成モデルとベイズ非パラメトリック報酬代理モデルを結合した、実証的に基づく再帰的アーキテクチャであるLarge Discovery Model(LDM)を紹介する。生成モデルは候補設計を提案・改良し、一方で代理モデルはその性能を予測して不確実性を定量化し、不確実性を考慮した価値を導出することで、候補の生成・改良・選択を導く。発見メモリと代理モデルは、新しい実験観測が得られるたびに継続的に更新される。我々は、ニューラルネットワーク学習、抗体設計、分子最適化を含む、異なる設計様式と目的関数にわたる3つのシナリオでLDMを評価した。LLMのみによる振り返りやこれらの領域にわたる従来の統計的探索と比較して、LDMは検証BPBの2.4倍の削減、結合エネルギーの18.2%の相対的減少、分子多目的性能の60%以上の相対的向上を達成した。これらの結果は、LDMが開かれた仮説空間における効果的な探索のための汎用発見エンジンとして機能し得ることを示唆している。
English
Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.