SLPO:通過代理策略擴展潛在推理
SLPO: Scaling Latent Reasoning via a Surrogate Policy
July 22, 2026
作者: Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li
cs.AI
摘要
基於可驗證獎勵的強化學習已成為引導顯式思維鏈推理器進行測試時擴展的主要方法。然而,這種擴展路徑在計算上仍然昂貴,因為每個中間步驟都必須解碼為語言標記。相反,隱式推理以連續向量承載中間計算,並且在遠更短的視野內已能匹配或超越顯式思維鏈。儘管具有此潛力,隱式推理器仍主要受侷限於模仿,而顯式思維鏈已通過結果獎勵強化學習超越了模仿。隱式軌跡缺乏可處理的每步似然性以及在固定思考預算下的自適應停止介面,因此結果獎勵無法引導隱式測試時擴展。我們引入替代隱式策略優化(SLPO),將結果獎勵強化學習帶入自迴歸隱式推理器:一種基於隱式轉換的實證替代策略密度,用於軌跡層級的信用分配;以及一個由正確性監督的停止頭,通過結果獎勵優化將其精煉為可變視野策略。在連續與軟性思考設定中,SLPO 提升了並行採樣下的 Pass@k 指標,並將更長的隱式計算分配給更困難的實例,同時保持更高的確定性準確率。
English
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across continuous and soft thinking settings, SLPO improves Pass@k under parallel sampling and allocates longer latent computation to harder instances with higher deterministic accuracy.