ChatPaper.aiChatPaper

오래되었지만 안정적인: 비동기 강화 학습 안정화를 위한 지연 적응형 신뢰 영역

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

July 21, 2026
저자: Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
cs.AI

초록

비동기 강화 학습은 롤아웃 생성과 최적화를 분리함으로써 처리량을 향상시키지만, 정책 지연, 엔진 지연, 그리고 전문가 혼합 라우팅으로 인해 복합적으로 발생하는 신선도 저하(staleness)는 불가피한 부산물이다. 신뢰 영역 관점에서 이러한 불일치는 매우 중요하다. 훈련-추론 발산은 유한 시간 범위에서의 근사 오차를 결정하는 반면, PPO 클리핑은 표본 추출된 외부 업데이트만을 선택적으로 차단하여 전체 정책 제약보다는 표본 기반 대리(sampled surrogate) 역할을 한다. 결과적으로, 신선도 저하가 심각한 업데이트는 낡은 롤아웃이 가장 중요한 비동기 방식에서 약하게 제어된다. 본 연구에서는 신선도 적응형 신뢰 영역(SAT)을 도입한다. 이 방법은 분리된 표본 로그 비율(detached sampled log-ratio)을 실용적인 신선도 저하 프록시로 사용하고, 신선도 기반 커널 스케일링을 통해 각 배치 내에서 높은 불일치를 보이는 꼬리 부분을 식별한 후, 명목상 PPO 구간에서 부호가 선택된 끝점만 수축시킨다. 이를 통해 일반적인 토큰에 대해서는 기준 행동을 유지하면서 새롭게 포착된 외부 대역에 대해서는 더 보수적인 업데이트를 적용한다. 우리는 국소적 구간 포함성(local interval containment)과 PPO 대비 점별 비관성(pointwise pessimism)을 증명하여, 적응형 규칙이 이질적인 신선도 저하 하에서 업데이트 기하학을 어떻게 재구성하는지 보여준다. 우리는 SGLang을 추론 엔진으로, Megatron을 훈련에 사용하여 Qwen3-30B-A3B-Base 기반의 분리된 비동기 RL 설정에서 SAT를 평가한다. 이 설정에서 SAT-GSPO w/ R3는 가장 우수한 AIME24 avg@8 성능을 기록하여 지연(lag) 1에서 35.83, 지연 8에서 34.79를 달성했으며, SAT-GSPO는 지연 1에서 34.17을 기록했다. 적응형 클리핑과 라우팅 재생(routing replay)은 각각 불일치 꼬리 부분과 라우팅 불일치를 대상으로 하는 상호 보완적인 안정화 장치로 작용한다. 전반적으로 클리핑 구간을 신선도 저하 이질성에 맞추는 것은 비동기 RL을 효과적으로 안정화한다.
English
Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint. As a result, high-staleness updates remain weakly controlled in the asynchronous regime where stale rollouts matter most. We introduce the Staleness-Adaptive Trust Region (SAT), which uses the detached sampled log-ratio as a practical staleness proxy, identifies high-mismatch tails within each batch via staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. This preserves baseline behavior on ordinary tokens while enforcing more conservative updates on newly intercepted outward bands. We prove local interval containment and pointwise pessimism relative to PPO, showing how the adaptive rule reshapes update geometry under heterogeneous staleness. We evaluate SAT in a decoupled asynchronous RL setup built on Qwen3-30B-A3B-Base, using SGLang as the inference engine and Megatron for training. In this setting, SAT-GSPO w/ R3 achieves the best observed AIME24 avg@8, reaching 35.83 at lag 1 and 34.79 at lag 8, while SAT-GSPO reaches 34.17 at lag 1. Adaptive clipping and routing replay act as complementary stabilizers targeting mismatch tails and routing inconsistency, respectively. Overall, aligning clip intervals with staleness heterogeneity effectively stabilizes asynchronous RL.