ChatPaper.aiChatPaper

효율적인 테스트 시점 스케일링을 위한 접두사 슬라이딩

Prefix Sliding for efficient test-time scaling

August 26, 2026
저자: Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, Ludwig Schmidt, Percy Liang, Jason Wei, Andrew Y. Ng, Luke Zettlemoyer, Yejin Choi, Mike Lewis
cs.AI

초록

테스트 시간 스케일링은 문제를 풀 때 언어 모델이 더 오래 추론하도록 하는 것처럼, 추가적인 테스트 시간 연산을 사용하여 성능을 향상시키는 기법이다. 모델이 전체 어텐션을 통해 추론 트레이스 전체를 메모리에 유지해야 하므로, 긴 사고를 필요로 하는 어려운 작업은 비용이 지나치게 커질 수 있다. 그러나 우리는 모델이 추론을 계속함에 따라 대부분의 중간 추론 토큰이 중요성을 잃는다는 것을 발견했다. 이는 그러한 토큰들을 유지하는 것이 비용 대비 가치가 있는지 의문을 제기한다. 이러한 통찰을 바탕으로, 우리는 추론 중에 프리픽스 또는 마지막 수천 개 토큰으로 구성된 윈도우에 속하지 않는 토큰을 폐기하는 Prefix Sliding을 제안한다. 프리픽스에는 모델이 활용할 수 있는 핵심 지침과 도구가 포함되고, 가장 최근 토큰들은 모델이 현재 수행 중인 추론을 나타낸다. 이 방식은 모델이 얼마나 오래 추론하든 전체 메모리 요구량을 일정 상한으로 제한하여, 효율적인 장기 테스트 시간 스케일링을 가능하게 한다. 학습 없이도 Prefix Sliding은 기존 모델의 속도를 3배 향상시키면서 성능을 유지할 수 있다. 강화 학습을 이용한 Prefix Sliding 훈련은 십만 개가 넘는 토큰에 이르는 추론 트레이스까지 스케일링을 가능하게 하여 더 나은 성능을 달성할 수 있다. 절제 실험 결과, Prefix Sliding은 중간 토큰을 요약하는 방식이나 기본 슬라이딩 윈도우보다 우수한 성능을 보였다. 우리의 코드는 https://github.com/Muennighoff/prefix-sliding 에 공개되어 있다.
English
Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive. However, we find most intermediate reasoning tokens lose importance as the model continues reasoning. This calls into question whether retaining them is worth the cost. Based on this insight, we propose Prefix Sliding, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens. The prefix has key instructions and tools available to the model, while the most recent tokens are the current reasoning the model is working on. This caps the total memory requirement regardless of how long the model reasons, allowing for efficient long-horizon test-time scaling. Without training, Prefix Sliding can make existing models 3x faster while maintaining performance. Training with Prefix Sliding using reinforcement learning can achieve better performance by enabling scaling to reasoning traces beyond a hundred thousand tokens. Ablations show Prefix Sliding outperforms summarizing intermediate tokens or vanilla sliding window. Our code is at https://github.com/Muennighoff/prefix-sliding