ChatPaper.aiChatPaper

オーディオ-ビデオモデルにおけるアテンショントライアングル

The Attention Triangle in Audio-Video Models

September 3, 2026
著者: Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes
cs.AI

要旨

音声-映像拡散モデルは、テキスト、音声、視覚コンテンツを協調させるためにクロスモーダル・アテンションに依存しているが、この同じ機構が、微妙かつ系統的な意味的漏洩を引き起こし得る。本稿では、テキスト、音声、映像の各ストリームを結ぶ三つのクロスアテンション・エッジから成る「アテンショントライアングル」を探索・分析し、生成中に意味情報がモダリティ間でどのようにルーティングされるかを調べる。分析の結果、音声-映像エッジに沿ったルーティングは双方向的であり、音声が映像生成に影響を及ぼす一方で、映像も音声生成に影響を及ぼすことが示された。このエッジは、モデルのパラメータに符号化されたバイアスによって形成され、漏洩の主要な要因となる。プロンプトと学習済み事前分布が衝突する場合、クロスモーダルな相互作用が意図された条件付けを上書きし、意味のルーティング先を、視覚的に典型的ではあるが誤った結果へと向け直し得る。これらの効果は、意味的アーティファクトが、単にアテンションが意図された対象を超えて拡散することから生じるのではなく、特定の経路に沿った構造化されたバイアス駆動型の相互作用から生じることを示唆している。この観点に基づき、我々は、意味がモダリティ間でどのように分布し接地されているかを明らかにするアテンション由来の信号を抽出し、それらを診断ツールとして用いることで、制御条件下での漏洩の分析と意図的な発生の両方を行う。これにより、クロスモーダルなルーティングの内部ダイナミクスを探り、個々の相互作用の役割を切り分けることが可能になる。さらに、これらの信号を利用して、より一貫したクロスモーダル整合を促す推論時の介入を導く。広範な実験は、本分析を裏付けるとともに、生成品質を維持しつつ意味的接地が改善されることを示している。
English
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.