外観ポインタ――拡散トランスフォーマーのマルチモーダル領域制御
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
July 21, 2026
著者: Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha
cs.AI
要旨
制御可能な画像生成は、創造的なプロフェッショナルにとって依然として課題であり、彼らはしばしば素材、オブジェクトのアイデンティティ、空間配置に対して、テキストプロンプトだけでは確実に実現できない精密な領域制御を必要とします。拡散トランスフォーマー(DiTs)は、テキストや画像に由来する異種トークンをネイティブに取り込むことができますが、これらのトークンが出力にどのように、どこで影響を与えるべきかを決定するメカニズムを欠いています。我々は「外観ポインタ(appearance pointers)」を導入します。これはコンパクトなトークンであり、テキストや画像の入力をユーザー指定のマスクと整合させることで、DiTを正しい空間位置における正しい外観の手がかりへと導きます。外観ポインタは領域対応ネットワークによって生成され、空間集約メカニズムを通じて洗練され、トークン負荷を大幅に増加させることなく、モデルが複数の領域記述を処理できるようにします。本手法は、基礎モデルをゼロから再トレーニングすることなく、DiTにおける局所的なマルチモーダル制御のための最初のモダリティ非依存インターフェースを導入します。さまざまな指標において、我々の単一モデルはモダリティ固有の最先端手法の性能に到達するか、それを上回り、生成的画像合成における精密で領域認識型のマルチモーダルガイダンスへのシンプルかつ拡張可能な道を提供します。
English
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.