SANA-Video 2.0: 効率的な動画生成のための注意残差を備えたハイブリッド線形注意機構
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
July 23, 2026
著者: Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie
cs.AI
要旨
本研究では、SANA-Video 2.0を紹介する。これは、統一アーキテクチャの下で5Bおよび14Bスケールで具現化されたハイブリッド動画拡散トランスフォーマーである。単一GPU上で最大720pの高品質動画を生成するよう設計されたSANA-Video 2.0は、完全ソフトマックス動画DiTの品質を維持しつつ、線形アテンションの長所である長いシーケンスへの好ましいスケーラビリティを保持する。全体に二次アテンションを避けるため、ハイブリッド線形-ソフトマックスアテンションでは、ゲート付き線形アテンションによるO(N)支配的な混合と、3:1の比率で周期的なゲート付きソフトマックスアンカーを組み合わせ、純粋な線形アテンションに欠けていたフルランクのトークン相互作用を回復する。これらの更新された表現を深さ方向に伝播させるために、ブロックアテンション残差(AttnRes)は、完了したブロック要約を後続の線形層にルーティングし、アンカー特徴の再利用を可能にし、深層の有効ランクを約12%向上させる。ゼロからのトレーニングにより、SANA-Video 2.0は、事前学習モデルを線形化するのではなく、完全なハイブリッドを直接学習し、低解像度でのプロキシ研究により、ソフトマックス割合25%が最適な品質と効率のトレードオフであることを確立する。40ステップのサンプリングで、SANA-Video 2.0は単一H100上で480p、13.2秒の処理でVBenchスコア84.30を達成し、はるかに大規模なソフトマックス動画DiTと遅延の低い部分で競争力を持つ。そのコンパイルされたDiTフォワードパスは、720p/60秒で同等の完全ソフトマックスベースラインより3.2倍高速であり、この差は動画の長さに応じて拡大する。さらに、フルスタックのSol-Engine最適化(カーネル融合、キャッシング、スパースアテンション)により、このハードウェアフレンドリーなバックボーンはさらに3.58倍高速化され、5Bパイプラインを720p/5秒で13.06秒に短縮し、1台のH100上でWan 2.2-A14Bよりも120倍高速になる。全体として、我々のハイブリッド設計は、コストを大幅に削減しながらソフトマックスレベルの表現力を回復し、スケーラブルで長時間・高解像度の動画生成を実現する。
English
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.