訓練具化學合理性感知的大型語言模型以進行單步逆合成
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
August 19, 2026
作者: Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
cs.AI
摘要
單步逆合成是電腦輔助合成規劃的核心組成部分,然而其本質上「一對多」的特性卻難以透過單一答案的評估與基準測試協議來充分呈現。為了解決此問題,我們引入Top-K提示作為一種穩健的訓練與推論範式,以更全面地捕捉多樣且合理的反應預測。我們彙編了CREED-CCV-2+USPTO-XL,這是一個包含約4,560萬條已驗證反應的超大規模資料集,用以訓練C3LM(化學約束一致語言模型)。透過將微調與基於ChemCensor及新穎性導向的獎勵機制相結合,我們的模型在OOD URSA-expert-2026基準上達到了最先進的效能。對反應唯一性的進一步分析顯示,大型語言模型與傳統模型探索了互補的反應空間,這為基於集成方法的逆合成系統提供了動機。整體而言,我們的結果確立了Top-K、具合理性感知的訓練作為未來穩健的基於大型語言模型之合成規劃的實用新方向。
English
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.