TurnOPD: ターン認識型オン・ポリシー蒸留による効率的な長期ホライズンエージェント訓練
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
July 7, 2026
著者: Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu, Jingjing Chen
cs.AI
要旨
オン方策蒸留(OPD)は、生徒自身の軌跡上でより強力な教師と一致させることで生徒方策を訓練し、言語エージェント訓練のための有望な枠組みを提供する。しかし、長期的なエージェントタスクへの応用は十分に探求されていない。我々は、バニラエージェントOPDにおける二つの主要な非効率性を特定する:(1) 全期間ロールアウトは、弱くノイズの多いKL監視を提供する末尾のターンにウォールクロックリソースを浪費することが多く、(2) 軌跡レベルのKL目的関数は損失のほとんどを浅いトークンに集中させ、初期の振る舞いが一致した後に深い意思決定ターンが十分に訓練されない。これらの課題に対処するため、我々はTurnOPDを提案する。これは長期的エージェントの効率的なオン方策蒸留のためのターンレベルの予算配分戦略である。TurnOPDは二つの予算制御器から構成される:プローブベースのターン統計量を用いてロールアウト長を決定する適応的ロールアウト深度予算配分と、KL重み付けをトークンレベルからターンバランス監視へと徐々に移行する段階的ターン正規化損失予算配分である。ALFWorld、WebShop、Multi-Hop Searchにおけるタスク特化型教師モデルを用いた実験は、TurnOPDが同等のウォールクロック訓練予算の下で優れた検証精度を達成し、精度-時間フロンティアをバニラOPDを超えて前進させることを示している。
English
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.