一段階逆合成のための化学的妥当性を考慮した大規模言語モデルの学習
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
August 19, 2026
著者: Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
cs.AI
要旨
一段階逆合成はコンピュータ支援合成計画の中核的な構成要素であるが、その本質的に一対多の性質は、単一解による評価・ベンチマークプロトコルでは十分に捉えられていない。この問題に対処するため、多様で妥当な反応予測をより良く捉える堅牢な学習・推論パラダイムとして、Top-Kプロンプティングを導入する。我々は、C3LM(化学的制約整合言語モデル、Chemistry Constraint-Consistent Language Model)の学習のために、約4560万件の検証済み反応からなる超大規模データセットCREED-CCV-2+USPTO-XLを構築した。ChemCensorに基づく報酬と新規性指向の報酬を用いた微調整を統合することで、本モデルはOOD URSA-expert-2026ベンチマークにおいて最先端の性能を達成する。反応のユニーク性に関するさらなる分析は、LLMと従来モデルが相補的な反応空間を探索することを示しており、アンサンブルベースの逆合成システムの動機付けとなる。全体として、本研究の成果は、Top-Kに基づく妥当性を考慮した学習を、将来の堅牢なLLMベース合成計画のための実用的な新たな方向性として確立するものである。
English
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.