ChatPaper.aiChatPaper

TILT:利用模型内在奖励提升扩散模型中的组合生成能力

TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

May 16, 2026
作者: Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury
cs.AI

摘要

近年来,强大的文本到图像生成模型的进步,使得开发测试时方法以修改采样轨迹,生成更符合复杂组合提示的图像变得日益重要。我们提出TILT,一种无需训练的框架,通过测试时奖励对齐实现组合文本到图像生成。我们将组合失败解释为联合概念分布与单一概念分布之间的重叠模式,并定义一种奖励,优先选择所有概念同时存在的样本。该奖励基于基础模型自身,无需任何外部监督或奖励模型。这产生了一个KL约束目标,具有封闭形式的倾斜目标分布以及用于扩散采样的原则性引导步骤。概念分布与上述奖励的相互作用自然导致了两种不同的引导策略,而平衡两者优势的混合方法则产生了更强的性能。在T2ICompBench的提示上进行的实验表明,与先前基线相比,我们的方法在保持图像质量的同时,改善了组合对齐。
English
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment. We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts are jointly present. This reward is intrinsic to the base model and does not require any external supervision or reward models. This yields a KL-constrained objective with a closed-form tilted target distribution and principled guiding steps for diffusion sampling. The interaction of concept distributions together with the above reward naturally leads to two different guidance strategies while a hybrid approach that balances their respective benefits produces stronger performance. Experiments on prompts from T2ICompBench show that our method improves compositional alignment while preserving image quality compared to previous baselines.