지지 집합 분할, 잔차 재구성: 비디오 생성 및 세계 모델을 위한 학습 불필요 희소 어텐션
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
August 19, 2026
저자: Pardis Taghavi, Reza Langari, Gaurav Pandey
cs.AI
초록
학습 없는 블록 희소 어텐션은 비디오 트랜스포머를 가속화할 수 있지만, 행 단위 어텐션 집중 자체만으로는 실행 가능한 희소 연산자를 규정하지 못한다. 블록 라우트를 공유하는 쿼리들은 지지 집합의 중첩이 낮을 수 있으며, 유지된 어텐션 질량만으로는 건너뛴 상호작용으로 인한 소프트맥스 후 오차를 결정할 수 없다. 우리는 분할 기하 구조가 통합된 지지 집합과 희소 출력으로부터 남은 잔차의 예측 가능성 모두에 영향을 미친다는 것을 보인다. 우리는 응답 결합 분할(Response-Coupled Partitioning)과 프로브 피팅 잔차 재구성(Probe-Fitted Residual Reconstruction)을 결합한 SparsePR을 제안한다. 샘플링된 쿼리-키 응답은 쌍을 이룬 K/V 그룹을 형성하며, 해당 그룹의 중심점은 공유 라우팅을 위한 쿼리-응답 좌표를 유도한다. 이후 소수의 정확한 쿼리 행은 프로브 잔차에서 관찰된 출력 부분공간 내에서 희소 출력으로부터 유도되는 호출별 아핀 보정을 산출한다. 네 가지 이질적인 비디오 생성 및 세계 모델 전반에 걸쳐 SparsePR은 어텐션 재구성 오차를 일관되게 줄인다. 절제 연구는 프로브 피팅이 이러한 감소의 대부분을 설명하며, 응답 결합 분할은 하드 드롭 오차를 낮추고 유한한 프로브 예산 하에서 재구성을 개선함을 보여준다. SparsePR은 22.0%~26.0%의 실현된 실행 쌍 밀도에서 생성 품질을 유지하면서 1.48배~2.61배의 종단 간 속도 향상을 달성한다. 프로젝트 페이지: https://pardistaghavi.github.io/SparsePR-website/
English
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post-softmax error from skipped interactions. We show that partition geometry affects both pooled support and the predictability of the remaining residual from the sparse output. We introduce SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. Sampled-query key responses form paired K/V groups, whose centroids induce query-response coordinates for shared routing. A small set of exact query rows then calibrates a call-specific affine correction from the sparse output within the output subspace observed in the probe residuals. Across four heterogeneous video generation and world models, SparsePR consistently reduces attention-reconstruction error. Ablations show that probe fitting accounts for most of this reduction, while response-coupled partitioning lowers hard-drop error and improves reconstruction under a finite probe budget. SparsePR preserves generation quality at 22.0-26.0% realized executed-pair density while achieving 1.48x-2.61x end-to-end speedups. Project page: https://pardistaghavi.github.io/SparsePR-website/