ChatPaper.aiChatPaper

외형 포인터 -- 확산 트랜스포머의 멀티모달 영역 제어

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

July 21, 2026
저자: Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha
cs.AI

초록

제어 가능한 이미지 생성은 창의적 전문가들에게 여전히 어려운 과제로 남아 있으며, 이들은 종종 텍스트 프롬프트만으로는 신뢰성 있게 달성할 수 없는 재료, 객체 정체성, 공간 배열에 대한 정밀한 지역적 제어를 필요로 한다. Diffusion Transformers(DiTs)는 텍스트와 이미지에서 비롯된 이종 토큰을 본래적으로 입력받을 수 있지만, 이러한 토큰이 출력에 어디에서 어떻게 영향을 미쳐야 하는지 결정하는 메커니즘이 부족하다. 우리는 외형 포인터(appearance pointers)를 도입하는데, 이는 텍스트 또는 이미지 입력을 사용자가 지정한 마스크와 정렬함으로써 DiT가 올바른 공간 위치에서 올바른 외형 단서를 찾도록 안내하는 간결한 토큰이다. 외형 포인터는 영역 대응 네트워크(region correspondence network)에 의해 생성되고, 공간 집계 메커니즘(spatial aggregation mechanism)을 통해 정제되어, 토큰 부하를 크게 증가시키지 않으면서 모델이 여러 지역적 설명을 처리할 수 있게 한다. 우리의 접근 방식은 기본 모델을 처음부터 재훈련하지 않고도 DiT에서 국소적 다중 모드 제어를 위한 최초의 모달리티 무관 인터페이스를 도입한다. 다양한 지표에 걸쳐, 우리의 단일 모델은 모달리티별 최신 방법의 성능에 도달하거나 이를 능가하며, 생성적 이미지 합성에서 정밀하고 영역 인식적인 다중 모드 안내를 위한 간단하고 확장 가능한 경로를 제공한다.
English
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.