RestoreKV: 공격적인 쿼리 무관 KV 캐시 축출 환경에서 전체 캐시 동작 복구
RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
August 2, 2026
저자: Changwoo Baek, Seungjun Shin, Kyeongbo Kong
cs.AI
초록
쿼리 무관 KV 캐시 축출은 컨텍스트를 한 번 압축하고 결과 캐시를 임의의 향후 쿼리에 재사용하지만, 제한된 예산에서는 성능이 붕괴될 수 있다. 기존 방법들은 주로 어떤 원본 KV 쌍을 유지할지 개선한다. 우리는 RestoreKV를 제안하며, 이는 동일한 총 KV 예산 내에서 학습된 복원을 통해 이러한 선택 기반 방식을 보완한다. 핵심 통찰은 축출로 인해 손실된 정보는 컨텍스트 특이적이지만, 그 간결한 보완 정보를 생성하는 메커니즘은 컨텍스트 간에 공유될 수 있다는 것이다. 컨텍스트 프리필 후, 소수의 복원 토큰이 단일 LoRA 적용 패스에서 전체 KV 캐시에 어텐션을 수행하여 컨텍스트 조건부 간결 복원 캐시를 생성한다. 기본 중요도 스코어러와 축출 규칙은 변경되지 않으며, 이후의 모든 쿼리와 디코딩에서는 어댑터가 비활성화된다. RestoreKV는 고정된 전체 캐시 모델로부터 매개변수 효율적 자기 증류를 통해 훈련되며, 매개변수의 0.4%만 최적화하고 작업별 튜닝을 요구하지 않는다. 네 개의 백본과 네 개의 장문 컨텍스트 벤치마크에서 RestoreKV는 압축으로 인한 성능 저하를 크게 줄인다. Qwen3-4B에서는 다섯 가지 기본 축출 방법에 걸쳐 예산이 일치하는 60개의 쌍 설정 중 59개를 개선한다. 5% 예산에서 RULER-4K의 KVzip 점수를 38.2에서 73.2로 향상시킨다. KVzip+에 적용하면 RestoreKV는 KVPress Benchmark에서 16배 압축 시 86.4의 RULER 정확도에 도달하며, 32K 컨텍스트 평가에서 0.5% 미만의 일회성 캐시 구축 오버헤드만 추가한다. 프로젝트 페이지는 https://paper.pnu-cvsp.com/RestoreKV/ 에서 확인할 수 있다.
English
Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained. We introduce RestoreKV, which complements this selection-based formulation with learned restoration under the same total KV budget. Our key insight is that, although the information lost through eviction is context-specific, the mechanism for generating its compact complement can be shared across contexts. After context prefill, a few restore tokens attend to the full KV cache in a single LoRA-adapted pass, generating a compact, context-conditioned restore cache. The base importance scorer and eviction rule remain unchanged, and the adapters are disabled for all subsequent queries and decoding. RestoreKV is trained through parameter-efficient self-distillation from the frozen full-cache model, optimizing only 0.4% of the parameters and requiring no task-specific tuning. Across four backbones and four long-context benchmarks, RestoreKV substantially reduces compression-induced degradation. On Qwen3-4B, it improves 59 of 60 paired, budget-matched settings across five base eviction methods; at a 5% budget, it raises KVzip from 38.2 to 73.2 on RULER-4K. Applied to KVzip+, RestoreKV reaches 86.4 RULER accuracy at 16times compression on the KVPress Benchmark, while adding less than 0.5% one-time cache-construction overhead in a 32K-context evaluation. Our project page is available at https://paper.pnu-cvsp.com/RestoreKV/