유클리드 클리핑을 넘어서: 리만 등거리 정책 최적화를 통한 LLM 강화학습의 탐색 붕괴 극복
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
July 11, 2026
저자: Zhicheng Cai, Xinyuan Guo, Hanlin Wu, Mingxuan Wang, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
cs.AI
초록
강화학습(RL)은 대규모 언어 모델(LLM)의 추론 능력을 향상시키는 지배적인 패러다임이 되었다. 그러나 PPO-Clip을 사용하는 RL 알고리즘은 본질적으로 탐색 붕괴(exploration collapse)에 제한된다. 이후의 연구들은 주로 경험적(heuristic) 접근에 머물러 있으며, PPO-Clip 실패의 근본 원인을 규명하지 못하고 있다. 본 연구는 PPO-Clip의 근본적 결함을 밝힌다. 즉, PPO-Clip은 정책 불일치(policy discrepancy)를 유클리드 거리(Euclidean metric)로 암묵적으로 측정하는데, 이는 정책 리만 다양체(policy Riemannian manifold)의 고유 기하와 이론적으로 일치하지 않는다. 이러한 기하학적 부정합(geometric mismatch)은 저확률 영역에서는 지나치게 보수적인 업데이트를, 고확률 영역에서는 공격적인 업데이트를 초래하여 궁극적으로 탐색을 붕괴시킨다. 이러한 기하학적 결함을 바로잡기 위해, 우리는 리만 다양체 상에서 등거리 정책 업데이트(isometric policy update)를 보장하여 탐색과 활용(exploitation)을 효과적으로 균형 맞추는 리만 등거리 정책 최적화(RIPO, Riemannian Isometric Policy Optimization)를 제안한다. 또한, RIPO가 최적화를 안정화하는 유리한 편향-분산 절충(bias-variance trade-off)을 달성함을 보인다. 광범위한 실험을 통해 RIPO가 7개의 경쟁 수준 벤치마크( AIME24에서 GRPO 대비 최대 60% 향상)에서 기존 LLM RL 알고리즘을 크게 능가함을 입증한다.
English
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identify the essential cause of PPO-Clip's failure. This work reveals the fundamental flaw of PPO-Clip: it implicitly measures policy discrepancy using Euclidean metric, which is theoretically inconsistent with the intrinsic geometry on the policy Riemannian manifold. This geometric mismatch results in overly conservative updates in low-probability regions while aggressive in high-probability regions, ultimately collapsing exploration. To correct this geometric flaw, we propose Riemannian Isometric Policy Optimization (RIPO), which guarantees isometric policy updates on the Riemannian manifold, effectively balancing exploration and exploitation. We further show that RIPO achieves a favorable bias-variance trade-off, which stabilizes optimization. Extensive experiments demonstrate that RIPO significantly surpasses existing LLM RL algorithms across seven competition-level benchmarks (up to 60% improvement over GRPO on AIME24).