ChatPaper.aiChatPaper

PCSD: 에이전트 강화 학습에서 자기 증류를 위한 지속적 일관성

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

August 3, 2026
저자: Chunji Lv, Yangguang Wei, Junlin Liu, Yang Gao, Ming Liu, Xinming Wang, Jinyang Wu, Guoren Wang, Changsheng Li
cs.AI

초록

대규모 언어 모델 에이전트는 복잡한 상호작용 과제에서 강력한 잠재력을 보여 주었지만, 길이가 긴 다중 턴 궤적이 단일한 결과 수준 신호만 받을 수 있기 때문에 강화 학습(RL)은 희소 보상에 의해 종종 저해된다. 온-폴리시 자기 증류(OPSD)는 특권 교사로부터 밀집된 토큰 수준의 감독을 제공하지만, 교사는 모든 위치에서 신뢰할 수 있는 것은 아니다. 기존 방법들은 일반적으로 노이즈에 민감할 수 있는 고립된 토큰 수준 불일치에 의존하거나, 위치 변동을 간과할 수 있는 공유 단계 수준 가중치를 할당한다. 우리는 교사 선호 신호의 국소적 지속성에서 토큰 수준 증류 가중치를 도출하는 지속 일관성 자기 증류(PCSD)를 제안한다. PCSD는 적응형 창과 지수 감쇠 집계를 결합하여 지속적인 상대적 교사 지지를 포착하고, 추세 인식 변조를 적용하여 국소적으로 감소하는 지지를 약화시키며, 시그모이드 게이팅을 통해 연속적인 가중치를 생성한다. 결과적인 목적 함수는 GRPO와 공동으로 최적화되어 밀집된 교사 안내와 희소한 환경 피드백을 결합한다. 추론 시 기술 없이도 PCSD는 두 백본 모두에서 모든 기준선 중 최고의 ALFWorld Overall 결과를 달성하여 GRPO를 15.6점 및 13.3점, SDAR를 6.2점 및 5.5점 초과하며, WebShop에서는 경쟁력을 유지하고 미보고 ALFWorld 분할에서 GRPO 대비 15.8점을 획득한다.
English
Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation (OPSD) provides dense token-level supervision from a privileged teacher, but the teacher may not be reliable at every position. Existing methods commonly rely on isolated token-level discrepancies, which can be sensitive to noise, or assign a shared step-level weight that may overlook positional variation. We propose Persistent Consistency Self-Distillation (PCSD), which derives token-level distillation weights from the local persistence of teacher-favoring signals. PCSD combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support, applies trend-aware modulation to attenuate locally declining support, and produces continuous weights through sigmoid gating. The resulting objective is jointly optimized with GRPO, combining dense teacher guidance with sparse environmental feedback. Without inference-time skills, PCSD achieves the best ALFWorld Overall results among all baselines on both backbones, exceeding GRPO by 15.6 and 13.3 points and SDAR by 6.2 and 5.5 points, while remaining competitive on WebShop and gaining 15.8 points over GRPO on unseen ALFWorld split.