ChatPaper.aiChatPaper

サポートを分割し、残差を再構成する:ビデオ生成とワールドモデルのための学習不要スパースアテンション

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

August 19, 2026
著者: Pardis Taghavi, Reza Langari, Gaurav Pandey
cs.AI

要旨

学習不要のブロックスパースアテンションはビデオトランスフォーマーを高速化できるが、行方向のアテンション集中それ自体は実行可能なスパース演算子を指定するものではない。ブロックルートを共有するクエリはサポートの重複が乏しい場合があり、一方で保持されるアテンション質量だけでは、スキップされた相互作用によるソフトマックス後誤差は決定されない。我々は、分割の幾何構造が、プールされたサポートと、スパース出力からの残差の予測可能性の両方に影響することを示す。我々は、応答結合分割とプローブ適合残差再構成を組み合わせたSparsePRを導入する。サンプリングされたクエリのキー応答は対になったK/Vグループを形成し、その重心が共有ルーティングのためのクエリ応答座標を誘導する。次に、少数の厳密なクエリ行が、プローブ残差で観測された出力部分空間内で、スパース出力からの呼び出し固有のアフィン補正を較正する。4つの異種のビデオ生成およびワールドモデルにわたり、SparsePRは一貫してアテンション再構成誤差を低減する。アブレーションは、この低減の大部分をプローブ適合が説明する一方、応答結合分割がハードドロップ誤差を低減し、有限のプローブ予算の下で再構成を改善することを示す。SparsePRは、22.0〜26.0%の実現された実行ペア密度で生成品質を維持しつつ、1.48倍〜2.61倍のエンドツーエンドの高速化を達成する。プロジェクトページ: https://pardistaghavi.github.io/SparsePR-website/
English
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post-softmax error from skipped interactions. We show that partition geometry affects both pooled support and the predictability of the remaining residual from the sparse output. We introduce SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. Sampled-query key responses form paired K/V groups, whose centroids induce query-response coordinates for shared routing. A small set of exact query rows then calibrates a call-specific affine correction from the sparse output within the output subspace observed in the probe residuals. Across four heterogeneous video generation and world models, SparsePR consistently reduces attention-reconstruction error. Ablations show that probe fitting accounts for most of this reduction, while response-coupled partitioning lowers hard-drop error and improves reconstruction under a finite probe budget. SparsePR preserves generation quality at 22.0-26.0% realized executed-pair density while achieving 1.48x-2.61x end-to-end speedups. Project page: https://pardistaghavi.github.io/SparsePR-website/