단일 단계 역합성을 위한 화학적 타당성 인식 대규모 언어 모델 학습
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
August 19, 2026
저자: Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
cs.AI
초록
단일 단계 역합성은 컴퓨터 지원 합성 계획의 핵심 구성 요소이지만, 그 본질적인 일대다 특성은 단일 정답 평가 및 벤치마킹 프로토콜로는 제대로 포착되지 않는다. 이러한 문제를 해결하기 위해, 우리는 다양하고 타당한 반응 예측을 더 잘 포착하기 위한 강건한 훈련 및 추론 패러다임으로서 Top-K 프롬프팅을 도입한다. 우리는 C3LM(화학 제약 일관 언어 모델)을 훈련하기 위해 약 4,560만 개의 검증된 반응으로 구성된 초대규모 데이터셋인 CREED-CCV-2+USPTO-XL을 구축한다. ChemCensor 기반 및 신규성 지향 보상을 통합한 미세 조정을 통해, 우리 모델은 OOD URSA-expert-2026 벤치마크에서 최첨단 성능을 달성한다. 반응 고유성에 대한 추가 분석은 LLM과 기존 모델이 상보적인 반응 공간을 탐색한다는 것을 보여주며, 이는 앙상블 기반 역합성 시스템을 위한 동기를 제공한다. 전반적으로, 우리의 결과는 Top-K 기반의 타당성 인지 훈련이 강건한 미래 LLM 기반 합성 계획을 위한 실용적인 새로운 방향임을 확립한다.
English
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.