ChatPaper.aiChatPaper

CONFLUX: RL後学習による3D胸部CT合成のための潜在拡散モデル

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training

July 3, 2026
著者: Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
cs.AI

要旨

制御可能な3次元医用画像生成モデルは、指定された臨床属性を持つボリュームを合成できるが、そのためには高い忠実度、本来の3次元性、そして要求された条件への忠実性を同時に満たすサンプルが必要となる。本論文では、胸部CT(コンピュータ断層撮影)のための潜在拡散モデルであるCONFLUXを提案する。3次元変分オートエンコーダが各ボリュームを圧縮し、整流フロートランスフォーマーが潜在空間で生成を行う。生成は、適応型層正規化を通じて構造化された放射線学的メタデータ(18種類の異常所見、性別、年齢、再構成カーネル)に条件付けられる。本モデルは、強力なボリュームベースライン(MAISIの74.6に対してFID 32.3)を達成し、臨床属性の直接的な制御を可能にする。この制御を強化するため、オンライン強化学習による事後学習段階(グループ相対ポリシー最適化)を追加し、生成された各ボリュームから分類器が要求された所見をどれだけ確実に復元できるかを報酬とする。別の独立した分類器による評価では、事後学習により実スキャンの信頼性に対する不足分の47%が解消された。本モデルと、幅広い臨床所見にわたる条件付けメタデータを備えた約20万件の合成胸部CTデータセットを公開する。
English
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning. We present CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space. Generation is conditioned on structured radiological metadata (18 abnormality findings, sex, age, and reconstruction kernel) through adaptive layer normalization. The model leads strong volumetric baselines on tri-planar Frechet distance (FID 32.3 vs. 74.6 for MAISI) while exposing direct control over clinical attributes. To strengthen that control we add an online reinforcement-learning post-training stage (group-relative policy optimization) that rewards how reliably a classifier recovers the requested findings from each generated volume. Judged by a separate, independent classifier, post-training removes 47% of the shortfall relative to real-scan reliability. We release the model and a ~200k synthetic chest-CT dataset with conditioning metadata spanning a wide variety of clinical findings.