ChatPaper.aiChatPaper

AgentOPSD: 에이전트 강화 학습을 위한 재귀적 자기 증류

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

August 6, 2026
저자: Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang
cs.AI

초록

검증 가능한 보상을 사용하는 강화 학습(RL)은 궤적 수준의 어드밴티지 추정치를 구성하지만, 장기 지평의 다중 턴 에이전트 작업에서 결과를 결정짓는 소수의 핵심 결정에 신용을 제대로 부여하지 못하는 경우가 많다. 최근 연구는 신용 할당을 위해 특권 자기 증류를 도입하여 더 조밀한 지도 신호를 제공하지만, 이러한 로컬 신호가 순차적 신용을 어떻게 표현해야 하는지는 여전히 불분명하다. 우리는 에이전트 강화 학습에서 턴 수준 신용 할당을 위한 비평가 없는 재귀적 방법인 AgentOPSD를 제안한다. AgentOPSD는 토큰 수준의 교사-학생 로그 확률 차이를 턴 수준 증거로 집계하고, 로그 오즈 공간에서 베이지안 신념 상태를 재귀적으로 갱신한다. 이를 통해 희소한 결과 감독을 턴 수준 신용 신호로 변환하는 원리적인 재가중 방식을 도출하며, 연속된 상태 간의 한계 신념 개정을 통해 핵심 턴을 식별한다. 이 방법은 표준 정책 최적화와 완전히 호환되며 추가 비평가도 추가 롤아웃도 요구하지 않는다. 우리는 Qwen2.5 모델을 두 가지 규모(3B 및 7B)로 사용하여 ALFWorld, WebShop, Search-QA에서 AgentOPSD를 평가한다. AgentOPSD는 GRPO 및 강력한 자기 증류 기준선보다 우수한 성능을 보이며, Qwen2.5-7B로 ALFWorld에서 89.1%의 성공률을 달성한다. 절제 연구는 이러한 성과가 턴 수준 집계와 이력 의존적 재귀 신념 갱신에 기인함을 보여준다.
English
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.