ChatPaper.aiChatPaper

SLPO: 대리 정책을 통한 잠재 추론 확장

SLPO: Scaling Latent Reasoning via a Surrogate Policy

July 22, 2026
저자: Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li
cs.AI

초록

검증 가능한 보상을 통한 강화 학습은 명시적 사고 사슬(Chain-of-Thought) 추론기에서 테스트 시간 확장을 유도하는 주요 방법이 되었다. 그러나 이 확장 경로는 모든 중간 단계를 언어 토큰으로 디코딩해야 하므로 계산 비용이 여전히 높다. 반면 잠재 추론은 중간 계산을 연속 벡터로 수행하며, 훨씬 짧은 수평선에서 명시적 CoT와 성능이 일치하거나 이를 능가한다. 이러한 가능성에도 불구하고 잠재 추론기는 대부분 모방 학습에 머물러 있는 반면, 명시적 CoT는 결과 보상 강화 학습을 통해 모방 학습을 넘어섰다. 잠재 궤적은 다루기 쉬운 단계별 우도와 고정된 사고 예산 하에서의 적응형 정지 인터페이스가 부족하기 때문에, 결과 보상으로는 잠재 테스트 시간 확장을 유도할 수 없다. 우리는 대리 잠재 정책 최적화(SLPO)를 도입하여 결과 보상 강화 학습을 자기회귀적 잠재 추론기에 적용한다. SLPO는 궤적 수준의 신용 할당을 위한 잠재 전이에 대한 경험적 대리 정책 밀도와, 결과 보상 최적화가 가변 수평 정책으로 정제하는 정확성 감독 정지 헤드로 구성된다. 연속 및 소프트 사고 설정 전반에서 SLPO는 병렬 샘플링 하에서 Pass@k를 개선하고, 더 어려운 사례에 더 긴 잠재 계산을 할당하여 더 높은 결정론적 정확도를 달성한다.
English
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across continuous and soft thinking settings, SLPO improves Pass@k under parallel sampling and allocates longer latent computation to harder instances with higher deterministic accuracy.