ChatPaper.aiChatPaper

TurnOPD: 효율적인 장기 지평 에이전트 훈련을 위한 온-정책 증류의 턴 인식화

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

July 7, 2026
저자: Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu, Jingjing Chen
cs.AI

초록

온-정책 증류(On-policy distillation, OPD)는 학생 정책이 자체 궤적에서 더 강력한 교사를 모방하도록 훈련함으로써, 언어 에이전트 훈련을 위한 유망한 프레임워크를 제공한다. 그러나 장기적 에이전트 과제에 대한 적용은 충분히 탐구되지 않았다. 우리는 기본 에이전트 OPD에서 두 가지 주요 비효율성을 확인한다: (1) 전체 시간 범위 롤아웃은 약하고 잡음이 많은 KL 감독을 제공하는 후반부 단계에 실제 시간 자원을 낭비하는 경우가 많으며, (2) 궤적 수준 KL 목표는 손실의 대부분을 표면적 토큰에 집중시켜 초기 행동이 정렬된 후에는 더 깊은 의사 결정 단계가 충분히 훈련되지 않게 한다. 이러한 문제를 해결하기 위해 우리는 장기적 에이전트의 효율적인 온-정책 증류를 위한 단계 수준 예산 전략인 TurnOPD를 제안한다. TurnOPD는 두 가지 예산 제어기로 구성된다: 탐침 기반 단계 통계를 사용하여 롤아웃 길이를 결정하는 적응형 롤아웃 깊이 예산 설정과, KL 가중치를 토큰 수준 감독에서 단계 균형 감독으로 점진적으로 전환하는 점진적 단계 정규화 손실 예산 설정이다. ALFWorld, WebShop 및 다중 홉 검색에서 과제 특화 교사 모델을 사용한 실험 결과, TurnOPD는 동일한 실제 시간 학습 예산 하에서 우수한 검증 정확도를 달성하고 기본 OPD를 넘어서는 정확도-시간 경계를 개선함을 보여준다.
English
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.