ChatPaper.aiChatPaper

外觀指針——擴散變換器的多模態區域控制

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

July 21, 2026
作者: Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha
cs.AI

摘要

可控图像生成对创意专业人士而言仍具挑战,他们常需对材质、物体身份及空间布局进行精确的区域控制,而仅靠文本提示无法可靠实现。扩散变换器(DiTs)能够原生处理来自文本和图像的异质令牌,但缺乏决定这些令牌在何处及如何影响输出的机制。我们提出了外观指针——紧凑型令牌,通过将文本或图像输入与用户指定的遮罩对齐,引导DiTs在正确的空间位置获取正确的外观线索。外观指针由区域对应网络生成,并通过空间聚合机制进行优化,使模型能够处理多个区域描述,同时不会显著增加令牌负载。我们的方法首次在无需从头重新训练基础模型的情况下,为DiT引入了模态无关的区域多模态控制接口。在一系列指标上,我们的单一模型达到或超越了特定模态的最先进方法,为生成式图像合成中精确、区域感知的多模态指导提供了一条简单且可扩展的路径。
English
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.