TILT:利用模型內在獎勵改善擴散模型中的組合生成
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
May 16, 2026
作者: Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury
cs.AI
摘要
近期,強大的文字到圖像生成模型取得了進展,使得開發測試時方法來修改取樣軌跡,以產生更忠於複雜組合提示的圖像變得日益重要。我們提出 TILT,這是一個透過測試時獎勵對齊來實現組合式文字到圖像生成的免訓練框架。我們將組合失敗解釋為聯合分布與單一概念分布之間的重疊模式,並定義一個獎勵,偏好所有概念共同存在的樣本。此獎勵是基底模型內在的,不需要任何外部監督或獎勵模型。這產生了一個具有封閉形式傾斜目標分布的 KL 約束目標,以及用於擴散取樣的原則性引導步驟。概念分布與上述獎勵的互動自然導致了兩種不同的引導策略,而一種平衡其各自優勢的混合方法則產生了更強的表現。在來自 T2ICompBench 的提示上進行的實驗顯示,與先前的基準相比,我們的方法在保持圖像品質的同時,改善了組合對齊。
English
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment. We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts are jointly present. This reward is intrinsic to the base model and does not require any external supervision or reward models. This yields a KL-constrained objective with a closed-form tilted target distribution and principled guiding steps for diffusion sampling. The interaction of concept distributions together with the above reward naturally leads to two different guidance strategies while a hybrid approach that balances their respective benefits produces stronger performance. Experiments on prompts from T2ICompBench show that our method improves compositional alignment while preserving image quality compared to previous baselines.