DrugGen 2: 疾患認識型言語モデルによる創薬の強化
DrugGen 2: A disease-aware language model for enhancing drug discovery
July 9, 2026
著者: Ali Motahharynia, Mohammadreza Ghaffarzadeh-Esfahani, Mahsa Sheikholeslami, Navid Mazrouei, Matin Irajpour, Yousof Gheisari, Hajar Sirous
cs.AI
要旨
現在の創薬における計算論的手法は、通常、特定の標的や一般的な分子特性に基づいて分子を生成することに焦点を当てており、疾患コンテキストが標的挙動や治療成績に与える影響をしばしば無視している。この課題に対処するため、我々は疾患オントロジーと標的タンパク質配列の両方を条件として低分子を設計する新規生成モデルDrugGen-2を提案する。DrugGen-2は、既承認薬とその疾患・標的を関連付けた厳選データセットに基づき、事前学習済みGPT-2モデルを教師ありファインチューニングと、それに続くグループ相対方策最適化(GRPO)による強化学習の2段階戦略を用いて微調整することにより開発された。このプロセスは、化学的妥当性、新規性、多様性、および高い予測結合親和性を最適化する報酬関数によって導かれた。糖尿病性腎症に関連する5つのタンパク質標的で評価したところ、DrugGen-2はベースラインモデル(DrugGPTおよびDrugGen)を有意に上回った。独自性の高い分子を生成する優れた能力を示し、既承認薬との構造的類似性が高く、全標的において改善された予測結合親和性を達成した。分子ドッキング解析によりこれらの知見はさらに支持され、強い結合ポテンシャルを有する候補リガンドが同定された。その中には、アンジオテンシン変換酵素に対する参照薬エナラプリルの結合親和性(-8.283)を上回る予測親和性(-9.917、-9.485、-9.367)を示す化合物も含まれていた。疾患特異的コンテキストを分子生成に統合することで、DrugGen-2はAI支援創薬を進展させ、疾患と分子標的の複雑な相互作用を考慮したデノボ設計やドラッグリポジショニングのための強力なツールを提供する。
English
Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences. DrugGen-2 was developed by fine-tuning a pre-trained GPT-2 model on a curated dataset of approved drugs linked to their diseases and targets, using a two-step strategy of supervised fine-tuning followed by reinforcement learning via group relative policy optimization (GRPO). This process was guided by reward functions optimizing for chemical validity, novelty, diversity, and high predicted binding affinity. When evaluated on five protein targets relevant to diabetic nephropathy, DrugGen-2 significantly outperformed baseline models (DrugGPT and DrugGen). It demonstrated a superior capacity to generate unique molecules, exhibited greater structural similarity to approved drugs, and achieved improved predicted binding affinities across all targets. Molecular docking analyses further supported these findings, identifying candidate ligands with strong binding potential, including compounds with predicted affinities (-9.917, -9.485, and -9.367) exceeding those of reference drugs such as enalapril for angiotensin-converting enzyme (-8.283). By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.