ChatPaper.aiChatPaper

多轮长程规划的物理学:通过单教师与多教师在策略智能体蒸馏从预训练到后训练

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

July 27, 2026
作者: Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao
cs.AI

摘要

多轮长程规划是基础模型智能体的关键能力,但其根本提升路径仍不明确。现有模型基于不可控且不透明的互联网数据进行训练,难以厘清规划能力如何获取、塑造与整合。为应对这一挑战,我们引入统一可控的多轮交互环境,实现精确控制,从而系统研究长程规划的三个阶段。(1)预训练阶段的规划能力获取。我们探究数据格式、分布与质量的影响:通过思维链状态转换建模构建显式世界模型,可增强长程泛化能力;原子技能不足以支撑组合泛化,而少量长程数据即可发挥效用;此外,次优轨迹会严重损害性能,原因在于错误在长程场景中会被逐级放大。(2)通过GRPO与OPD进行后训练的规划能力塑造。利用互信息,我们将通用规划模式与任务特有规划知识区分开来。针对规划模式,我们识别出后训练的三个应用区域:不必要、有效和不可支持。在低质量与长程场景下,OPD的有效区域比GRPO更广,因其提供更一致的更新方向。针对规划知识,若教师具有不同知识,从其蒸馏未见过程可能损害学生已有的世界模型,而未能充分建立新知识。(3)通过MOPD后训练实现规划能力整合。研究表明,多教师同策略蒸馏(MOPD)通过在不同环境中收敛至共享规划模式来整合能力:兼容模式可实现跨环境泛化,部分共享模式支持持续学习,而完全冲突模式则引发严重干扰。
English
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student's prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.