TILT: 모델 내재적 보상을 통한 확산 모델의 구성적 생성 개선
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
May 16, 2026
저자: Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury
cs.AI
초록
최근 강력한 텍스트-이미지 생성 모델의 발전으로 인해, 복잡한 구성적 프롬프트에 더 충실한 이미지를 생성하기 위해 샘플링 궤적을 수정하는 테스트 시점 방법을 개발하는 것이 점점 더 중요해지고 있습니다. 우리는 테스트 시점 보상 정렬을 통한 구성적 텍스트-이미지 생성을 위한 학습 불필요 프레임워크인 TILT를 제시합니다. 우리는 구성적 실패를 결합 분포와 단일 개념 분포 간의 중복 모드로 해석하고, 모든 개념이 함께 존재하는 샘플에 유리한 보상을 정의합니다. 이 보상은 기본 모델에 고유하며 외부 감독이나 보상 모델을 필요로 하지 않습니다. 이는 닫힌 형태의 기울어진 목표 분포와 확산 샘플링을 위한 원칙적인 유도 단계를 갖춘 KL 제약 목적 함수를 생성합니다. 개념 분포와 위 보상의 상호작용은 자연스럽게 두 가지 다른 유도 전략을 이끌어내며, 각각의 이점을 균형 잡는 하이브리드 접근법이 더 강력한 성능을 보입니다. T2ICompBench의 프롬프트에 대한 실험은 우리의 방법이 이전 기준선과 비교하여 이미지 품질을 유지하면서 구성적 정렬을 개선함을 보여줍니다.
English
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment. We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts are jointly present. This reward is intrinsic to the base model and does not require any external supervision or reward models. This yields a KL-constrained objective with a closed-form tilted target distribution and principled guiding steps for diffusion sampling. The interaction of concept distributions together with the above reward naturally leads to two different guidance strategies while a hybrid approach that balances their respective benefits produces stronger performance. Experiments on prompts from T2ICompBench show that our method improves compositional alignment while preserving image quality compared to previous baselines.