ChatPaper.aiChatPaper

SANA-Video 2.0:基於注意力殘差的混合線性注意力機制實現高效視頻生成

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

July 23, 2026
作者: Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie
cs.AI

摘要

我們介紹SANA-Video 2.0,這是一種混合式影片擴散Transformer,在統一架構下分別以5B和14B兩種規模實例化。此模型專為在單一GPU上生成高達720p的高品質影片而設計,在品質上可與全softmax影片DiT匹敵,同時保留了線性注意力在長序列擴展上的優勢。為避免整體使用二次注意力,混合線性-軟最大值注意力(Hybrid Linear-Softmax Attention)將門控線性注意力(用於O(N)主導的混合)與週期性門控軟最大值錨點以3:1的比例結合,恢復了純線性注意力所缺乏的全秩令牌交互。為了在深度上傳播這些更新後的表示,區塊注意力殘差(AttnRes)將完成的區塊摘要路由至後續線性層,實現錨點特徵重用,並將深層有效秩提升約12%。透過從頭訓練,SANA-Video 2.0直接學習完整的混合結構,而非將預訓練模型線性化,同時透過降低解析度的代理研究確立25%的softmax為最佳品質與效率權衡。採用40步取樣時,SANA-Video 2.0在單一H100上以480p解析度、13.2秒內達到VBench分數84.30,與規模更大的softmax影片DiT相比,在延遲大幅降低的情況下仍具競爭力。其編譯後的DiT前向傳遞在720p/60秒時比匹配的全softmax基線快3.2倍,且差距隨影片時長擴大。此外,全棧Sol-Engine優化(核心融合、快取與稀疏注意力)進一步將此硬體友善的骨幹網路加速3.58倍,使5B管線在720p/5秒時達到13.06秒,在單一H100上比Wan 2.2-A14B快120倍。總體而言,我們的混合設計以顯著降低的成本恢復了softmax級別的表達力,實現了可擴展的長時長、高解析度影片生成。
English
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.