ChatPaper.aiChatPaper

FVAttn: 비디오 생성을 위한 런타임 부하 분산을 적용한 적응형 희소 어텐션

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

July 17, 2026
저자: Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen, Hao Liu, Mohan Zhang, Chen Li, Ziyang Ma, Jing Lyu, Jiangsu Du
cs.AI

초록

비디오 확산 트랜스포머는 긴 시공간 시퀀스를 처리하므로, 자기 주의(self-attention)가 고해상도 비디오 생성의 주요 병목 현상이 됩니다. 훈련 없는 희소 주의는 이 비용을 줄이지만, 적응형 Top-p 라우팅은 다중 GPU 시퀀스 병렬 처리에서 헤드별 작업 부하를 불균등하게 만듭니다. 그 결과 작업 부하 이질성은 희소 주의를 랭크 수준의 지연(straggler) 문제로 만듭니다. 본 논문에서는 다중 GPU 시퀀스 병렬 처리에서 적응형 희소 주의의 분산 실행 효율성을 개선하는 훈련 없는 희소 주의 시스템인 을 제시합니다. 은 희소 라우팅 프론트엔드로 Top-p 라우팅, Top-k 안전 하한 및 비디오 인식 블록 구성을 사용한 후, 런타임에 구체화된 마스크를 수정합니다. 런타임 부하 균등화는 P2P 통신을 통해 소수의 무거운 헤드를 마이그레이션하여 현재 중요 경로를 단축합니다. 여유 인식 희소 증강은 잔여 비중요 랭크 여유를 추가적인 고가치 블록으로 채우며, 오버랩은 기존 계산 뒤에 스케줄링 및 마이그레이션 오버헤드를 숨깁니다. 단계 증류된 Wan2.2 I2V에서 은 평균 부하 불균형을 1.34에서 1.08로 줄이고 FlashAttention 대비 4.41배의 주의 속도 향상을 제공하며, 경쟁력 있는 비디오 품질로 2.02~2.11배의 DiT 추론 속도 향상을 달성합니다.
English
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present , a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. uses Top-p routing, a Top-k safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, reduces average load imbalance from 1.34 to 1.08 and delivers a 4.41times attention speedup over FlashAttention, while achieving a 2.02--2.11times DiT inference speedup with competitive video quality.