ChatPaper.aiChatPaper

KVpop -- 採用預測性在線剪枝的鍵值快取壓縮

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

July 6, 2026
作者: Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl, David Stap, Thomas Schmied, Sebastian Böck, Günter Klambauer, Sepp Hochreiter
cs.AI

摘要

鍵值(KV)緩存增長是自迴歸解碼中的一個主要瓶頸,因為記憶體與頻寬會隨上下文長度線性擴展。現有的KV逐出方法通常依賴靜態啟發式或代理分數,這些方法難以追蹤未來標記的效用,一旦相關性發生變化,就會導致脆弱的逐出。為了解決此問題,我們引入了KVpop,它透過直接監督保留或丟棄決策,學習一種固定預算的KV逐出策略。該評分器針對一種新穎的未來注意力目標進行訓練,該目標可在不實體化密集注意力圖的情況下高效計算。我們進一步引入了一種基於延遲記憶的評分器,這在學習型逐出方法中獨樹一幟,它將評分推遲固定步數,以利用近未來上下文。在AIME和HMMT數學推理任務中,KVpop在Qwen3-4B模型上,以75%的KV緩存壓縮率保留了98%的全注意力效能,以88%的壓縮率保留了97%的效能,始終優於既有的逐出基線。Qwen3-8B模型表現出更強的結果,接近完全教師模型的效能。這些結果表明,使用未來注意力信號監督逐出,能夠在降低記憶體成本的同時保持品質。
English
Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, which poorly track future token utility and cause brittle eviction as relevance shifts. To address this, we introduce KVpop, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision. The scorer is trained against a novel future-attention target, computed efficiently without materializing dense attention maps. We further introduce a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context. On AIME and HMMT mathematical reasoning, KVpop retains 98% of full-attention performance on Qwen3-4B at 75% KV cache compression and 97% at 88% compression, consistently outperforming established eviction baselines. Qwen3-8B shows even stronger results, reaching near-full teacher performance. These results show that supervising eviction with future-attention signals cuts memory costs while maintaining quality.