FVAttn: 動画生成のための実行時負荷分散を伴う適応的スパース注意機構
FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
July 17, 2026
著者: Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen, Hao Liu, Mohan Zhang, Chen Li, Ziyang Ma, Jing Lyu, Jiangsu Du
cs.AI
要旨
ビデオ拡散トランスフォーマーは長い時空間シーケンスを処理するため、高解像度動画生成において自己注意が主要なボトルネックとなる。訓練不要のスパース注意はこのコストを削減するが、適応型Top-pルーティングはマルチGPUシーケンス並列処理下でヘッドごとに不均一なワークロードを生み出す。このワークロードの不均一性により、スパース注意はランクレベルのストラグラー問題へと変質する。本論文では、マルチGPUシーケンス並列処理下での適応型スパース注意の分散実行効率を向上させる、訓練不要のスパース注意システムを提案する。本システムは、スパースルーティングのフロントエンドとしてTop-pルーティング、Top-kセーフティフロア、および動画認識型ブロック編成を採用し、実行時にマテリアライズドマスクを修復する。ランタイム負荷分散では、P2P通信により少数の重いヘッドを移行して現在のクリティカルパスを短縮する。スラック認識型スパース拡張では、残りの非クリティカルランクスラックを高価値ブロックで追加的に埋める一方、オーバーラップによりスケジューリングと移行のオーバーヘッドを既存の計算の背後に隠蔽する。ステップ蒸留されたWan2.2 I2Vにおいて、本システムは平均負荷不均衡を1.34から1.08に低減し、FlashAttention比で4.41倍の注意高速化を実現するとともに、競争力のある動画品質を維持しつつ2.02~2.11倍のDiT推論高速化を達成する。
English
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present , a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. uses Top-p routing, a Top-k safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, reduces average load imbalance from 1.34 to 1.08 and delivers a 4.41times attention speedup over FlashAttention, while achieving a 2.02--2.11times DiT inference speedup with competitive video quality.