ChatPaper.aiChatPaper

音視頻模型中的注意力三角

The Attention Triangle in Audio-Video Models

September 3, 2026
作者: Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes
cs.AI

摘要

音訊-視訊擴散模型依賴跨模態注意力來協調文字、聲音與視覺內容;然而,正是此機制可能引入細微且系統性的語義洩漏。為了研究這些模型,我們透過探測並分析「注意力三角形」——由連接文字流、音訊流與視訊流的三條跨模態注意力邊所構成——來檢視語義資訊在生成過程中如何被跨模態路由。 我們的分析揭示,沿著音訊-視訊邊的路由是雙向的:音訊可以影響視訊生成,而視訊也可以影響音訊生成。這條邊由模型參數中編碼的偏誤所塑造,並成為洩漏的主要促成因素:當提示與學習到的先驗相牴觸時,跨模態互動可能凌駕於原本的條件設定之上,將語義重新路由至視覺上典型但錯誤的結果。這些效應顯示,語義偽影的產生不僅來自注意力擴散至其原有目標之外,更源於沿特定路徑、由偏誤驅動的結構化互動。 立足於此觀點,我們提取注意力導出的訊號,以揭露語義如何在各模態之間分布與扎根,並將其用作診斷工具,在受控條件下既能分析洩漏,也能刻意誘發洩漏。這使我們得以探測跨模態路由的內部動態,並分離出個別互動所扮演的角色。我們進一步利用這些訊號引導推論時的干預措施,以促成更一致的跨模態對齊。大量實驗支持我們的分析,並證明在維持生成品質的同時,語義扎根有所改善。
English
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.