ChatPaper.aiChatPaper

並非所有Token都應獲得同等信用:面向長思維鏈推理的反事實敏感性信用重新分配

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

July 30, 2026
作者: Qiangqiang He, Zhongheng Wu, ZiJian Wang
cs.AI

摘要

帶有可驗證獎勵的強化學習(RLVR)對於改善大型語言模型中的長思維鏈推理至關重要。無評論家方法(如 GRPO)將回應層級獎勵轉換為優勢,並均勻地廣播至所有 token,忽略了它們對最終結果的不平等貢獻。相反地,在策略自蒸餾(OPSD)透過最小化非特權策略與特權自教師之間的前向 KL 散度,提供密集的分佈式監督,其隱含假設是:由此產生的似然偏移編碼了可靠的、與答案對齊的資訊。我們透過固定每個取樣軌跡,並在兩種相反的結果條件下重新評分來檢驗此前提:一種條件斷言正確,另一種斷言不正確。大多數受影響的 token 在兩種條件下朝相同方向偏移,很少出現符號反轉,且所誘發的最佳化訊號有顯著重疊。大的偏移也集中在高度可替代的表面形式 token 上,而攜帶問題特定推理內容的 token 則較不敏感。這些發現表明,特權偏移無法提供可靠的、與答案對齊的方向,而其幅度主要反映反事實敏感性,而非 token 層級的學習價值。基於這些觀察,我們提出反事實敏感性信用重新分配(CSCR),這是 GRPO 的一個簡單擴展,它減少對高度敏感 token 的信用,並重新正規化 token 層級優勢,以保留原始的信用預算和驗證器決定的方向。在長思維鏈數學推理基準上,CSCR 在相同數量的策略更新下持續優於 GRPO 基線。有針對性的消融實驗進一步證實了我們的診斷:特權誘導的方向不可靠,適度的降權最有效,而更強烈的調制會使最佳化不穩定。
English
Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional supervision by minimizing the forward KL divergence between an unprivileged policy and a privileged self-teacher, implicitly assuming that the resulting likelihood shifts encode reliable answer-aligned information. We test this premise by fixing each sampled trajectory and re-scoring it under two opposing outcome conditions, one asserting correctness and the other incorrectness. Most affected tokens shift in the same direction under both conditions, with few sign reversals and substantial overlap in the induced optimization signals. Large shifts also concentrate on highly substitutable surface-form tokens, whereas tokens carrying problem-specific reasoning content are less sensitive. These findings show that privileged shifts fail to provide reliable answer-aligned directions, while their magnitudes primarily reflect counterfactual sensitivity rather than token-level learning value. Based on these observations, we propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve both the original credit budget and verifier-determined direction. On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms GRPO baseline with the same number of policy updates. Targeted ablations further corroborate our diagnosis: privilege-induced directions are unreliable, moderate downweighting is most effective, and stronger modulation destabilizes optimization.