CONFLUX:一種用於三維胸部CT合成且結合強化學習後訓練的潛在擴散模型
CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
July 3, 2026
作者: Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
cs.AI
摘要
可控生成式三维医学影像模型能够合成具有特定临床属性的体素数据,但这要求样本同时具备高保真度、原生三维特性以及对条件约束的忠实还原。我们提出CONFLUX——一种面向胸部计算机断层扫描(CT)的潜在扩散模型:通过三维变分自编码器压缩每个体素数据,并采用整流流变换器在潜在空间中进行生成。生成过程通过自适应层归一化,以结构化放射学元数据(18种异常发现、性别、年龄和重建内核)为条件进行调控。该模型在三平面弗雷歇距离(FID)指标上显著超越强体积基线模型(CONFLUX得分32.3,MAISI得分74.6),同时实现了对临床属性的直接控制。为增强这种控制力,我们引入在线强化学习后训练阶段(组相对策略优化),通过奖励机制提升分类器从生成体积中准确恢复指定发现的可靠性。经独立分类器评估,后训练过程消除了与真实扫描可靠性之间的47%差距。我们开源该模型及约20万例合成胸部CT数据集,其条件元数据涵盖广泛临床发现类型。
English
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning. We present CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space. Generation is conditioned on structured radiological metadata (18 abnormality findings, sex, age, and reconstruction kernel) through adaptive layer normalization. The model leads strong volumetric baselines on tri-planar Frechet distance (FID 32.3 vs. 74.6 for MAISI) while exposing direct control over clinical attributes. To strengthen that control we add an online reinforcement-learning post-training stage (group-relative policy optimization) that rewards how reliably a classifier recovers the requested findings from each generated volume. Judged by a separate, independent classifier, post-training removes 47% of the shortfall relative to real-scan reliability. We release the model and a ~200k synthetic chest-CT dataset with conditioning metadata spanning a wide variety of clinical findings.