멀티턴 장기 계획의 물리학: 단일 및 다중 교사의 온-폴리시 에이전틱 증류를 통한 사전 학습에서 사후 학습까지
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
July 27, 2026
저자: Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao
cs.AI
초록
다중 턴 장기 계획은 기반 모델 에이전트에 있어 중요하지만, 이를 근본적으로 개선하는 방법은 여전히 불분명하다. 기존 모델은 통제 불가능하고 불투명한 인터넷 데이터로 학습되어, 계획 능력이 어떻게 습득되고 형성되며 통합되는지 파악하기 어렵다. 이러한 문제를 해결하기 위해, 우리는 정밀한 제어가 가능한 통합적이고 통제된 다중 턴 환경을 도입한다. 이를 통해 세 단계에 걸쳐 장기 계획을 체계적으로 연구할 수 있다. (1) 사전 학습 중 계획 능력 습득. 데이터 형식, 분포, 품질을 연구한다. CoT 상태 전이 모델링을 통한 명시적 세계 모델 구축은 더 강력한 장기 일반화를 제공한다. 원자적 기술만으로는 구성적 일반화에 충분하지 않지만, 소량의 장기 데이터는 효과적이다. 또한, 차선의 궤적은 오류가 장기간에 걸쳐 증폭되기 때문에 성능을 심각하게 저하시킨다. (2) GRPO와 OPD 사후 학습을 통한 계획 능력 형성. 상호 정보를 통해 일반 계획 패턴과 작업별 계획 지식을 구분한다. 계획 패턴에 대해 사후 학습의 세 가지 적용 영역(불필요, 효과적, 지원 불가)을 식별한다. OPD는 저품질 및 장기 설정에서 GRPO보다 더 넓은 효과 영역을 가지는데, 이는 보다 일관된 업데이트 방향을 제공하기 때문이다. 계획 지식의 경우, 다른 지식을 가진 교사로부터 보지 못한 절차를 증류하면 새로운 지식을 완전히 확립하지 못하면서 학생의 기존 세계 모델을 손상시킬 수 있다. (3) MOPD 사후 학습을 통한 계획 능력 통합. 다중 교사 온-정책 증류(MOPD)가 환경 간 공유된 계획 패턴으로 수렴함으로써 능력을 통합함을 보여준다. 호환 가능한 패턴은 환경 간 일반화를 가능하게 하고, 부분적으로 공유된 패턴은 지속적 학습을 지원하며, 완전히 충돌하는 패턴은 심각한 간섭을 유발한다.
English
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student's prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.