TILT: モデル内在的報酬による拡散モデルの構成的生成の改善
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
May 16, 2026
著者: Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury
cs.AI
要旨
近年、強力なテキスト・画像生成モデルの進展に伴い、複雑な構成的プロンプトにより忠実な画像を生成するために、サンプリング軌道を修正するテスト時手法の重要性が高まっている。本稿では、テスト時の報酬アライメントを通じて構成的テキスト・画像生成を行う訓練不要のフレームワークTILTを提案する。構成的失敗を、結合概念分布と単一概念分布間のオーバーラップモードとして解釈し、すべての概念が共に存在するサンプルを優先する報酬を定義する。この報酬はベースモデルに内在するものであり、外部の教師信号や報酬モデルを必要としない。これにより、閉形式の傾斜目標分布と拡散サンプリングのための原理的なガイダンスステップを備えたKL制約付き目的関数が得られる。概念分布と上記の報酬の相互作用は、自然に2つの異なるガイダンス戦略をもたらし、それぞれの利点をバランスするハイブリッドアプローチがより強力な性能を発揮する。T2ICompBenchのプロンプトを用いた実験により、本手法が従来のベースラインと比較して画質を維持しつつ、構成的アライメントを改善することを示す。
English
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment. We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts are jointly present. This reward is intrinsic to the base model and does not require any external supervision or reward models. This yields a KL-constrained objective with a closed-form tilted target distribution and principled guiding steps for diffusion sampling. The interaction of concept distributions together with the above reward naturally leads to two different guidance strategies while a hybrid approach that balances their respective benefits produces stronger performance. Experiments on prompts from T2ICompBench show that our method improves compositional alignment while preserving image quality compared to previous baselines.