ChatPaper.aiChatPaper

多輪長程規劃的物理機制:透過單教師與多教師的同策略智能體蒸餾從預訓練到後訓練

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

July 27, 2026
作者: Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao
cs.AI

摘要

多輪長程規劃對於基礎模型代理至關重要,然而如何從根本上提升此能力仍不明確。現有模型訓練於不可控且不透明的網路資料,難以釐清規劃能力的獲取、塑造與整合過程。為解決此挑戰,我們引入一個統一且可控的多輪環境,得以精確調控,並系統性地研究長程規劃的三個階段:(1) 預訓練階段的規劃能力獲取。我們探討資料格式、分佈與品質。透過思維鏈狀態轉換建模進行明確的世界模型建構,能產生更強的長程泛化能力。僅靠原子技能不足以實現組合泛化,而少量的長程資料便能奏效。此外,次優軌跡會嚴重損害效能,原因在於誤差在長程中會累積放大。(2) 經由GRPO與OPD後訓練的規劃能力塑造。透過互信息,我們區分通用規劃模式與任務特定的規劃知識。針對規劃模式,我們辨識出後訓練的三種應用區域:不必要、有效與無法支撐。在低品質與長程設定下,OPD的有效區域較GRPO更廣,因其提供更一致的更新方向。針對規劃知識,從具不同知識的教師模型中蒸餾未見過的程序,可能損害學生既有的世界模型,同時未能完整建立新知識。(3) 經由MOPD後訓練的規劃能力整合。我們證明多教師同策略蒸餾(MOPD)能透過收斂至環境間共享的規劃模式來整合能力。相容的模式能實現跨環境泛化,部分共享的模式支援持續學習,而完全衝突的模式則導致嚴重干擾。
English
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student's prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.