ChatPaper.aiChatPaper

KVpop -- 키-값 캐시 압축을 위한 예측적 온라인 가지치기

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

July 6, 2026
저자: Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl, David Stap, Thomas Schmied, Sebastian Böck, Günter Klambauer, Sepp Hochreiter
cs.AI

초록

키-값(KV) 캐시 증가는 자동회귀 디코딩의 주요 병목 현상으로, 메모리와 대역폭이 컨텍스트 길이에 비례하여 선형적으로 증가합니다. 기존의 KV 제거 방법은 종종 정적 휴리스틱 또는 대리 점수에 의존하는데, 이는 미래 토큰의 유용성을 제대로 추적하지 못하고 관련성이 변화함에 따라 취약한 제거를 초래합니다. 이러한 문제를 해결하기 위해, 우리는 유지 또는 제거 결정을 직접 지도하여 고정 예산의 KV 제거 정책을 학습하는 KVpop을 소개합니다. 점수화기는 새로운 미래-어텐션 목표에 대해 학습되며, 이는 조밀한 어텐션 맵을 구체화하지 않고 효율적으로 계산됩니다. 또한, 학습 기반 제거 방법 중에서 유일하게, 근미래 컨텍스트를 활용하기 위해 점수화를 고정된 스텝 수만큼 지연시키는 지연된 메모리 기반 점수화기를 도입합니다. AIME 및 HMMT 수학적 추론 과제에서 KVpop은 Qwen3-4B 모델의 75% KV 캐시 압축 시 전체 어텐션 성능의 98%를 유지하고, 88% 압축 시 97%를 유지하며, 기존의 제거 기준 방법들을 일관되게 능가합니다. Qwen3-8B 모델은 더욱 강력한 결과를 보여주며, 거의 완전한 교사 성능에 도달합니다. 이러한 결과는 미래-어텐션 신호로 제거를 지도함으로써 품질을 유지하면서 메모리 비용을 절감할 수 있음을 보여줍니다.
English
Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, which poorly track future token utility and cause brittle eviction as relevance shifts. To address this, we introduce KVpop, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision. The scorer is trained against a novel future-attention target, computed efficiently without materializing dense attention maps. We further introduce a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context. On AIME and HMMT mathematical reasoning, KVpop retains 98% of full-attention performance on Qwen3-4B at 75% KV cache compression and 97% at 88% compression, consistently outperforming established eviction baselines. Qwen3-8B shows even stronger results, reaching near-full teacher performance. These results show that supervising eviction with future-attention signals cuts memory costs while maintaining quality.