잠재 온-폴리시 자기 증류
Latent On-Policy Self-Distillation
August 13, 2026
저자: Guibin Zhang, Jiayang Lyu, Ran Sun, Xinlei Yu, Haoyu Zhao, Qibing Ren, Shuicheng Yan
cs.AI
초록
에이전트가 경험으로부터 학습하고 이를 정책에 내재화하는 것은 자기 진화 AI의 핵심 문제가 되었다. OPSD(On-Policy Self-Distillation)는 특권적 자기 교사를 활용하여 학생 자신의 궤적에 대해 밀집 지도를 제공함으로써 효과적인 경로를 제시한다. 그러나 기존 방법들은 여전히 설계자가 지정한 특권적 인공물(예: 답변, 피드백, 스킬, 궤적)에 크게 의존하며, 이는 지속적 자기 개선에 필요한 종단 간 학습 가능성과 확장성을 제한한다. 본 연구에서는 LOPD(Latent On-Policy Self-Distillation)를 소개한다. LOPD는 새롭게 규정된 형태의 특권적 맥락을 가진 또 다른 수작업 OPSD 변형을 제안하는 대신, 교사의 특권적 맥락 자체를 경험으로부터 종단 간 학습 가능하게 만든다. 기술적으로 LOPD는 관련 경험을 검색하여 이를 연속적인 잠재 토큰으로 구성하고, 이 토큰이 자기 교사를 조건화한다. 학생은 과제와 상호작용 이력으로부터 궤적을 생성하며, 방문한 모든 접두사(prefix)에서 토큰 수준의 밀집 지도를 받는다. 또한, 잠재 맥락 학습을 안정화하고 조절하기 위해 특권적 마진 목적 함수를 도입한다. 실험적으로 LOPD는 (I) 에이전트형 도구 사용과 코드 생성 모두에서 RLVR 및 OPSD, SDPO, Skill-SD를 포함한 대표적인 OPSD 방법들을 능가하는 강력한 성능을 보여주며, (II) 롤아웃 예산의 30% 미만만으로 GRPO 및 Skill-SD를 능가하는 높은 학습 효율을 달성한다. 절제 연구는 특권적 맥락을 학습 가능하게 만드는 것이 이러한 성과를 실현하는 데 필수적이라는 직접적인 증거를 추가로 제공한다. 종합적으로, 이러한 결과는 LOPD를 에이전트 진화를 위한 보다 확장 가능하고 자기 주도적인 패러다임으로 나아가는 한 단계로 자리매김한다.
English
Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged artifacts (e.g., answers, feedback, skills, or trajectories), limiting the end-to-end learnability and scalability required for continual self-improvement. In this work, we introduce Latent On-Policy Self-Distillation (LOPD), which, rather than proposing another hand-crafted OPSD variant with a newly prescribed form of privileged context, makes the teacher's privileged context itself learnable end-to-end from experience. Technically, LOPD retrieves relevant experiences and composes them into continuous latent tokens that condition a self-teacher, while the student generates trajectories from the task and interaction history and receives dense token-level supervision at every visited prefix. We further introduce a privileged-margin objective to stabilize and regulate the learning of latent context. Empirically, LOPD demonstrates (I) strong performance, outperforming RLVR and representative OPSD methods including OPSD, SDPO, and Skill-SD across both agentic tool use and code generation; and (II) high learning efficiency, surpassing GRPO and Skill-SD with less than 30% of their rollout budget. Ablation studies further provide direct evidence that making privileged context learnable is necessary for realizing these gains. Together, these results position LOPD as a step toward a more scalable and self-directed paradigm for agent evolution.