ChatPaper.aiChatPaper

남은 것을 학습하라, 숙달된 것을 학습하지 말라: 다중 보상 정책 최적화를 위한 포화 인지 어드밴티지 재가중

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

August 17, 2026
저자: Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, Yijiang Li
cs.AI

초록

그룹 상대적 이점을 사용한 강화 학습(RL)은 언어 모델 추론기의 후속 학습에 있어 사실상의 표준이 되었다. 그러나 여러 보상 목표를 최적화할 때, 기존 방법들은 일반적으로 그룹별 표준화 전에 보상 벡터를 고정된 가중 합으로 스칼라화한다. 우리는 이러한 설계가 두 가지 근본적인 문제를 초래함을 보여준다: 서로 다른 보상 프로파일을 가진 롤아웃들이 동일한 이점을 받을 수 있으며, 모든 목표가 현재의 포화(saturation) 수준과 무관하게 고정된 상대 가중치로 최적화된다는 점이다. 결과적으로 훈련은 더 큰 잔여 개선 여지가 있는 목표에 집중하는 대신, 이미 해결된 목표에 그래디언트 예산을 계속 할당한다. 우리는 다중 보상 정책 최적화를 위한 포화 인지 이점 재가중(SA-MRPO)을 소개한다. 이 방법은 각 보상 목표를 독립적으로 표준화하고, 목표 포화도의 배치 수준 추정치에 따라 해당 목표의 기여도를 적응적으로 할인한다. 이는 이미 충분히 충족된 목표에 대한 성능을 경험적으로 유지하면서, 최적화 노력을 덜 최적화된 목표로 동적으로 재배분한다. 우리는 또한 포화 인지 재가중이 단순히 업데이트의 크기를 재조정하는 것에 그치지 않고, 업데이트의 부호를 반전시킬 수 있음을 보여준다. 두 개 및 세 개의 보상 목표 조합을 사용한 수학적 추론 전반에서, SA-MRPO는 15개의 벤치마크 비교 중 12개에서 GDPO보다 더 어려운 정확성 목표를 개선하며, AIME24에서 최대 5%의 향상을 보인다. 적응형 추론에서는 다섯 가지 벤치마크 모두에서 정확도를 개선하며, 평균 3.8%, AMC23에서 최대 9.2%의 향상을 보인다. 코딩 벤치마크에서는 통과율을 최대 2.3% 향상시키며, 모든 설정에서 더 쉬운 목표를 이미 충족된 수준 근처로 유지한다.
English
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. As a result, training continues to allocate gradient budget to already-solved objectives instead of focusing on those with greater remaining headroom. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This dynamically reallocates optimization effort toward under-optimized objectives while empirically maintaining performance on those that are already well satisfied. We further show that saturation-aware reweighting can reverse the sign of an update, rather than merely rescale its magnitude. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to 5% on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by 3.8% on average and up to 9.2 % on AMC23, and on coding benchmarks it improves pass rate by up to 2.3%, while in all settings maintaining the easier objectives near their already satisfied levels.