マルチターン長期計画の物理学:事前学習から事後学習への単一・複数教師オン方策エージェンティック蒸留
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
July 27, 2026
著者: Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao
cs.AI
要旨
マルチターン長期計画は基盤モデルエージェントにとって重要であるが、その根本的な改善方法は依然として不明確である。既存のモデルは制御不能で不透明なインターネットデータで訓練されており、計画能力がどのように獲得され、形成され、統合されるかを特定することは困難である。この課題に対処するため、我々は統一された制御可能なマルチターン環境を導入し、精密な制御を可能にする。これにより、以下の三つの段階にわたって長期計画を体系的に研究できる。(1) 事前学習における計画能力の獲得。データ形式、分布、品質を研究する。CoT状態遷移モデリングによる明示的な世界モデルの構築は、より強力な長期汎化をもたらす。原子的スキルだけでは構成汎化に不十分であり、少量の長期計画データが有効である。さらに、準最適な軌道は、長期的な誤差の増幅により性能を著しく損なう。(2) GRPOおよびOPDによる事後学習における計画能力の形成。相互情報量を通じて、一般的な計画パターンとタスク固有の計画知識を区別する。計画パターンについて、事後学習の三つの適用領域(不要、有効、非対応)を特定する。OPDは、低品質および長期間の設定においてGRPOよりも広い有効領域を持ち、より一貫した更新方向を提供する。計画知識については、異なる知識を持つ教師から未観測の手順を蒸留すると、新しい知識を完全に確立せずに生徒の既存の世界モデルを損なう可能性がある。(3) MOPD事後学習による計画能力の統合。マルチ教師オンポリシー蒸留(MOPD)は、環境間で共有される計画パターンに収束することで能力を統合することを示す。互換性のあるパターンは環境間汎化を可能にし、部分的に共有されたパターンは継続学習を支援するが、完全に競合するパターンは深刻な干渉を引き起こす。
English
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student's prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.