ChatPaper.aiChatPaper

올바른 계층적 희소 어텐션: 무한 컨텍스트 모델링을 위하여

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

July 3, 2026
저자: Xiang Hu, Xinyu Wei, Hao Gu, Minshen Zhang, Tian Liang, Huayang Li, Lei Zhu, Yan Wang, Sirui Han, Yushi Bai, Kewei Tu, Haitao Mi, Leo Liang
cs.AI

초록

현대 대규모 언어 모델(LLM)을 긴 컨텍스트로 확장하는 것은 이차 계산 비용과 밀집 어텐션의 열악한 길이 외삽으로 인해 제한된다. 청크 단위 희소 어텐션은 유망한 대안을 제공하지만, 기존 모든 방법은 부정확한 청크 선택으로 인해 완전 어텐션에 미치지 못한다. 본 논문에서는 언어 모델링(LM) 손실 하에서 종단간 청크 선택을 학습하는 청크 단위 희소 어텐션 메커니즘인 계층적 랜드마크 희소(HiLS) 어텐션을 제안한다. HiLS는 어텐션을 계층적으로 분해한다: 각 쿼리는 검색된 각 청크와 독립적으로 어텐션을 수행하여 청크별 정보를 추출하고, 결과 출력은 청크 검색 점수에 따라 융합된다. HiLS는 검색 점수를 순방향 어텐션 계산에 통합함으로써 LM 손실로 직접 최적화하여 종단간 검색 학습과 본질적 희소 학습을 가능하게 한다. 실험 결과는 HiLS-어텐션이 도메인 내 컨텍스트 길이에서 완전 어텐션에 필적하거나 일부 경우 더 나은 성능을 달성함을 보여준다. 동시에 HiLS-어텐션은 훈련 컨텍스트 길이의 64배 이상을 90% 검색 정확도로 외삽하며, 이는 완전 어텐션을 훨씬 능가한다. 또한, 기존 완전 어텐션 모델은 경량 지속 사전학습을 통해 HiLS-어텐션으로 변환될 수 있으며, 도메인 내 성능을 유지하면서 초장기 컨텍스트 외삽 능력을 획득한다. HiLS-어텐션은 희소 KV 접근 및 계산과 함께 일반적인 효율성-성능 트레이드오프를 깨뜨리며, 완전 어텐션 기반 모델보다 일반적인 장기 컨텍스트 작업에서 더 효율적이고 더 효과적인 장기 컨텍스트 LLM을 가능하게 한다.
English
Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of dense attention. Chunk-wise sparse attention offers a promising alternative, but all existing methods fall short of full attention because of their inaccurate chunk selection. We propose Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling (LM) loss. HiLS factorizes attention hierarchically: each query performs attention independently with each retrieved chunk to extract chunk-specific information, and the resulting outputs are fused according to chunk retrieval scores. By incorporating retrieval scores into the forward attention computation, HiLS optimizes them directly with the LM loss, enabling end-to-end retrieval learning and native sparse training. Experimental results show that HiLS-Attention achieves performance comparable to, and in some cases better than, full attention at in-domain context lengths. Meanwhile, HiLS-Attention extrapolates more than 64times the training context length with 90% retrieval accuracy, far beyond full attention. Moreover, existing full-attention models can be converted to HiLS-Attention with lightweight continued pretraining, preserving in-domain performance while acquiring ultra-long-context extrapolation. Together with its sparse KV access and computation, HiLS-Attention breaks the usual efficiency-performance trade-off, enabling long-context LLMs that are both more efficient and more effective on general long-context tasks than their full-attention counterparts.