SANA-Video 2.0:混合线性注意力与注意力残差实现高效视频生成
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
July 23, 2026
作者: Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie
cs.AI
摘要
我们提出SANA-Video 2.0,这是一种在统一架构下以5B和14B规模实例化的混合视频扩散Transformer。该模型专为在单GPU上生成长达720p的高质量视频而设计,在质量上可与全softmax视频DiT相媲美,同时保留了线性注意力在长序列处理上的优势。为了避免全程使用二次注意力,混合线性-softmax注意力机制以3:1的比例将门控线性注意力(实现O(N)主导的混合)与周期性门控softmax锚点相结合,恢复了纯线性注意力所缺失的全秩token交互。为在深度上传播这些更新的表示,块注意力残差(AttnRes)将完成的块摘要路由至后续线性层,实现了锚点特征复用,并将深层有效秩提升约12%。通过从头训练,SANA-Video 2.0直接学习完整的混合机制,而非对预训练模型进行线性化处理,并通过降分辨率代理研究确定了25% softmax为质量与效率的最佳折衷点。采用40步采样时,SANA-Video 2.0在单块H100上以480p分辨率、13.2秒内达到VBench评分84.30,在延迟远低于更大规模softmax视频DiT的情况下保持竞争力。其编译后的DiT前向传播在720p/60秒场景下比匹配的全softmax基线快3.2倍,且该差距随视频时长增加而扩大。此外,全栈Sol-Engine优化(内核融合、缓存和稀疏注意力)将此硬件友好型主干网络进一步加速3.58倍,使得5B流水线在720p/5秒场景下仅需13.06秒,比单块H100上的Wan 2.2-A14B快120倍。总体而言,我们的混合设计以大幅降低的成本恢复了softmax级别的表现力,从而解锁了可扩展的长时、高分辨率视频生成能力。
English
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.