ChatPaper.aiChatPaper

BeaconKV: 대규모 추론 모델의 효율적인 추론을 위한 비콘 쿼리 기반 키-값 캐시 압축

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

September 4, 2026
저자: Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi
cs.AI

초록

대규모 추론 모델(LRM)은 확장된 사고 연쇄(CoT) 생성을 통해 뛰어난 문제 해결 능력을 달성하지만, 그 결과 발생하는 키-값(KV) 캐시는 시퀀스 길이에 따라 선형적으로 증가하여 심각한 메모리 병목을 초래하며, 긴 추론 추적에서는 종종 GPU 용량을 초과한다. 기존 KV 캐시 압축 방법들은 미래 토큰 중요도를 추정하기 위해 최근 쿼리에 의존하며, 이러한 쿼리가 미래 어텐션 패턴에 대한 신뢰할 수 있는 대리 지표로 작용한다고 암묵적으로 가정한다. 우리는 이러한 가정이 장기적 추론에서 실패함을 입증한다. 특정 디코딩 단계에서는 추론 추적 초기에 수립된 문제 해결 계획과 같이 멀리 떨어진 이전 문맥을 다시 어텐션하는 사고 재방문 토큰(TRT)이 생성된다. 체계적 분석을 통해 우리는 TRT에 대응하는 쿼리들이 임베딩 공간에서 소수의 유사도 그룹으로 군집화된다는 점을 발견한다. 이러한 통찰에 기반하여, 우리는 전체 쿼리 이력을 저장하지 않고 어떤 KV 쌍이 재방문될지 예측하기 위해 각 전역 쿼리 클러스터에 대한 간결한 대표자인 비콘 쿼리를 유지하는, 학습이 필요 없는 KV 캐시 압축 방법인 BeaconKV를 제안한다. 네 가지 오픈소스 LRM과 다양한 추론 벤치마크 전반에서 BeaconKV는 일반적으로 기존 압축 방법들을 능가하며, 전체 캐시 정확도를 거의 유지하면서 최대 5.8배의 메모리 절감을 달성하고 처리량을 4.3배 이상 향상시킨다.
English
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) that re-attend to distant previous context, such as task-solving plans formulated early in the trace. Through systematic analysis, we discover that queries corresponding to the TRT cluster into a small number of similarity groups in the embedding space. Based on this insight, we propose BeaconKV, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history. Across four open-source LRMs and diverse reasoning benchmarks, BeaconKV generally outperforms existing compression methods, achieving up to 5.8times memory reduction while nearly preserving full cache accuracy and improving throughput by over 4.3times.