ChatPaper.aiChatPaper

뮤온(Muon)은 에이전트 강화 학습에서 언제 도움이 되는가?

When Does Muon Help Agentic Reinforcement Learning?

July 17, 2026
저자: Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
cs.AI

초록

Muon은 대규모 사전 학습에서 AdamW와 경쟁력을 보이지만, 강화학습(RL) 사후 학습에서의 가치는 아직 명확하지 않습니다. 우리는 Qwen2.5-0.5B-Instruct를 사용한 ALFWorld에서 일치된 단일 시드 비교를 통해 희소 보상 에이전트 RL에 바닐라 Muon을 적용하여 AdamW와 대조했습니다. Group-in-Group Policy Optimization (GiGPO)에서 Muon을 은닉 가중치 행렬에만 적용하면 최종 윈도우 검증 성공률이 0.290에서 0.546으로 (+88%) 상승했습니다. 반면 고율 AdamW 제어군은 업데이트 이후 성공률이 전혀 유지되지 않았습니다. 이러한 효과는 이점 추정기와 학습률에 따라 달라집니다. 3e-5에서 Muon은 GRPO의 성공률을 0.161에서 0.268로 향상시킨 반면, GraphGPO의 후기 윈도우 격차는 포화점 근처에서 좁혀집니다. 1e-5에서 GraphGPO Muon은 0.901에 도달하고 정규화된 검증 AUC를 0.399에서 0.556으로 높였으며, 성공률 0.5와 0.75에 각각 30회 및 60회 더 일찍 도달했습니다. 이러한 탐색적 결과는 Muon이 에이전트 RL에 이점을 제공할 수 있음을 보여주며, 정책 최적화기, 이점 추정기, 학습률을 함께 연구할 동기를 부여합니다. 다중 시드 및 교차 작업 검증은 여전히 과제로 남아 있습니다.
English
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The effect depends on the advantage estimator and learning rate. At 3e-5, Muon improves GRPO from 0.161 to 0.268, whereas GraphGPO's late-window gap narrows near saturation. At 1e-5, GraphGPO Muon reaches 0.901, raises normalized validation AUC from 0.399 to 0.556, and reaches 0.5 and 0.75 success 30 and 60 updates earlier, respectively. These exploratory results show that Muon can benefit agentic RL and motivate studying the policy optimizer, advantage estimator, and learning rate jointly. Multi-seed and cross-task validation remain open.