ChatPaper.aiChatPaper

群體熵控制策略優化

Group Entropy-Controlled Policy Optimization

July 18, 2026
作者: Guangran Cheng, Chengqi Lyu, Songyang Gao, Wenwei Zhang, Kai Chen
cs.AI

摘要

熵控制已成為大型語言模型(LLM)強化學習(RL)中的有效工具,有助於在對齊過程中平衡探索-利用取捨。此類強化學習範式通常針對異質任務的混合進行,這些任務在同一策略下會誘發不同的熵區間,使得全局或詞元層級的熵調節不足以應對相應的異質探索需求。這種異質性進一步導致GRPO風格的歸一化優勢產生熵依賴偏差,使得不同提示組之間的優勢訊號在統計上不具可比性。為解決此問題,我們提出群組熵控制策略優化(GEPO),這是對GRPO的輕量級擴展,利用從現有分組樣本中估計的群組熵來進行基於熵條件的非對稱優勢塑形。GEPO在低熵群組中衰減正向優勢以減少過度利用,並在高熵群組中衰減負向優勢以保留探索,且其自適應閾值源自歷史熵統計。在橫跨數學、物理、科學、程式碼生成與指令遵循等十三個基準測試的兩個基礎模型上進行的廣泛實驗表明,GEPO持續優於GRPO及近期基於熵控制的方法,在訓練過程中提供平衡的跨任務改進,同時保留任務特定的探索水準。
English
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of heterogeneous tasks, which induce distinct entropy regimes under the same policy, making global or token-level entropy regulation insufficient to corresponding heterogeneous needs of exploration. This heterogeneity further makes GRPO-style normalized advantages induce an entropy-dependent bias, making advantage signals across prompt groups statistically non-comparable. To address this issue, we propose Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that uses group entropy, estimated from existing grouped samples to perform entropy-conditioned asymmetric advantage shaping. GEPO attenuates positive advantages in low-entropy groups to reduce over-exploitation, and negative advantages in high-entropy groups to preserve exploration, with adaptive thresholds derived from historical entropy statistics. Extensive experiments on two base models across thirteen benchmarks spanning mathematics, physics, science, code generation, and instruction following show that GEPO consistently outperforms GRPO and recent entropy-controlled methods, delivering balanced cross-task improvements while preserving task-specific exploration levels throughout training.