ChatPaper.aiChatPaper

FVAttn:具有运行时负载均衡的自適應稀疏注意力用於視頻生成

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

July 17, 2026
作者: Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen, Hao Liu, Mohan Zhang, Chen Li, Ziyang Ma, Jing Lyu, Jiangsu Du
cs.AI

摘要

視訊擴散變換器處理長時空序列,使得自注意力成為高解析度視訊生成的主要瓶頸。訓練免稀疏注意力降低了此成本,但自適應 Top-p 路由在多 GPU 序列並行下會產生不均勻的每頭計算負載。由此產生的負載異質性將稀疏注意力轉變為秩層級的拖後腿問題。我們提出 ,這是一套訓練免稀疏注意力系統,可在多 GPU 序列並行下提升自適應稀疏注意力的分散式執行效率。 採用 Top-p 路由、Top-k 安全下限及視訊感知區塊組織作為稀疏路由前端,並在運行時修復具體化的遮罩。運行時負載平衡透過點對點通訊遷移少量重頭來縮短當前關鍵路徑。寬鬆感知稀疏增強則以額外高價值區塊填補殘餘的非關鍵秩空閒,同時重疊運算隱藏排程與遷移的開銷。在步蒸餾的 Wan2.2 I2V 上, 將平均負載不平衡從 1.34 降至 1.08,並相較於 FlashAttention 實現 4.41 倍的注意力加速,同時在具競爭力的視訊品質下達到 2.02 至 2.11 倍的 DiT 推論加速。
English
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present , a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. uses Top-p routing, a Top-k safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, reduces average load imbalance from 1.34 to 1.08 and delivers a 4.41times attention speedup over FlashAttention, while achieving a 2.02--2.11times DiT inference speedup with competitive video quality.