ChatPaper.aiChatPaper

CRISP: 절벽 인식 입력 적응형 희소 프리필링 및 구조 질량 기반 라우팅

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

September 1, 2026
저자: Huu Huy Nguyen, Chien Van Nguyen, Franck Dernoncourt, Ryan A. Rossi, Linh Ngo Van, Jieyang Chen, Thien Huu Nguyen
cs.AI

초록

장문 컨텍스트 LLM 추론의 어텐션 프리필(prefilling) 단계는 이차적으로 확장되므로, 셀프 어텐션은 심각한 계산 병목 현상을 유발한다. 기존의 희소 어텐션(sparse attention) 방법은 고정된 패턴이나 오프라인 프로파일링을 통해 이를 완화하지만, 입력 의존적인 어텐션 구조에 적응할 수 있는 유연성이 부족하다. 최근의 동적 방법은 헤드를 희소 패턴으로 실시간 라우팅하여 이를 해결하지만, 간접적인 라우팅 프록시를 사용하여 오버헤드가 발생하고, 소프트맥스 이후 질량 계층 구조를 간과하는 예산 할당 메커니즘에 의존한다. 우리는 CRISP(Cliff-awaRe Input-adaptive Sparse Prefilling)를 제시하여, 이러한 동적 라우팅 패러다임의 두 가지 구조적 과제를 식별하고 해결한다. 첫째, 라우팅 결정이 프록시 어텐션 맵의 구조에서 직접 읽혀질 수 있음을 보여준다. 우리는 Jensen-Shannon Divergence(JSD) 라우팅을 C_struct라는 구조적 프록시로 대체한다. C_struct는 Vertical-Slash 호환 위치에서 질량을 측정하며, 풀링된 행렬 곱셈과 이후의 KL 발산 오버헤드를 모두 제거하면서 JSD의 라우팅 결정을 재현한다. 둘째, 소프트맥스 이후 질량 절벽(mass cliff)을 공식화하고, 엄격한 누적 커버리지 임계값이 긴 컨텍스트에서 O(n)의 배경 잡음을 축적함을 이론적으로 증명한다. CRISP는 잡음 바닥(noise floor)에 기반한 싱크 인지 임계값(sink-aware threshold)을 통해 이를 탐색한다. 실증적으로, 두 모델 계열에 걸쳐 InfiniteBench, RULER 및 LongBench에서 CRISP는 전반적으로 가장 강력한 희소 방법이며, 검색 집약적 벤치마크에서 정확한 dense 어텐션과 동등하거나 이를 능가하여 검색 작업에서 기준선 대비 최대 +28.0 pp의 성능 회복을 달성하고, 512k 토큰에서 최대 5.30배의 어텐션 속도 향상을 달성한다. 이는 주로 선택 과정에서 구조적 무결성을 보존하면서 O(n) 잡음을 제거함으로써 실현된다.
English
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe computational bottleneck. Traditional sparse attention methods mitigate this through fixed patterns or offline profiling, but lack the flexibility to adapt to input-dependent attention structure. Recent dynamic methods address this by routing heads to sparse patterns in real-time, but rely on indirect routing proxies with overhead and budget allocation mechanisms that overlook the post-softmax mass hierarchy. We present CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling), which identifies and addresses two structural challenges in this dynamic routing paradigm. First, we show that the routing decision can be read directly off the structure of the proxy attention map. We replace the Jensen-Shannon Divergence (JSD) routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead. Second, we formalize the post-softmax mass cliff and demonstrate theoretically that strictly cumulative coverage thresholds accumulate O(n) background noise at long contexts. CRISP navigates this via a sink-aware threshold grounded in the noise floor. Empirically, across InfiniteBench, RULER and LongBench on two model families, CRISP is the strongest sparse method overall and matches or exceeds exact dense attention on retrieval-heavy benchmarks, recovering up to +28.0 pp on retrieval tasks over baselines and achieving up to a 5.30x attention speedup at 512k tokens, driven primarily by our O(n) noise elimination during selection while preserving structural integrity.