SimpleOPD: 장문맥 추론을 위한 간단한 토크나이저 독립적 온-정책 증류
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
August 14, 2026
저자: Haonan He, Haodi Lei, Yun Luo, Haoran Zhang, Shunkai Zhang, Yizhuo Li, Shengji Tang, Zhilin Wang, Runzhe Zhan, Lei Bai, Ganqu Cui, Fangchen Yu, Yafu Li, Peng Ye, Ning Ding, Yu Cheng
cs.AI
초록
온-정책 증류(OPD)는 강한 교사 모델의 추론 능력을 전이하는 유망한 방법이지만, 이를 장문맥 추론 교사 모델과 단문맥 학생 모델에 적용하는 것은 토크나이저 불일치, 교사-학생 분포 불일치, 응답 길이 폭발, 훈련 불안정성 등의 실질적인 문제를 야기한다. 본 연구에서는 장문맥 추론 모델인 SU-01의 증명 추론 능력을 단문맥 학생 모델로 전이함으로써 이러한 설정을 연구한다. 토크나이저 차이를 처리하기 위해 공유 텍스트 공간에서 OPD를 수행하고, 학생 및 교사 토크나이저에서 동일한 텍스트 범위를 차지하는 토큰만 정렬한다. 과도한 생성 길이와 빈번한 절단 문제를 완화하기 위해 학생 참조 KL 손실을 도입하고 </think> 및 <|im_end|>와 같은 특수 종료 토큰의 어드밴티지를 마스킹한다. 이 전략은 학생이 초기 정책에서 과도하게 이탈하지 않도록 제약하여 교사-학생 분포 불일치 문제를 완화하고 안정적인 길이 성장을 촉진한다. Qwen3, Qwen3.5, Intern-S2, GLM-4.7, Gemma-4를 포함한 동일 계열 및 이종 계열 학생 모델에 대한 실험은 수학적 추론, 특히 자연어 수학 증명에서 일관된 성능 향상을 보여준다. 특히 Intern-S2-Preview는 ProofBench에서 21.2점 향상되어 55.2점을 달성하며 Gemini-2.5-Pro를 능가한다. 또한 HLE 및 HiPhO와 같은 과학 벤치마크에서도 개선되어, OPD가 수학 훈련 영역을 넘어 일반화되는 추론 능력을 전이함을 시사한다.
English
On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mismatch, teacher-student distribution mismatch, response length explosion, and training instability. In this work, we study this setting by transferring proof-reasoning capabilities from the long-context reasoning model SU-01 to short-context student models. To handle tokenizer differences, we perform OPD in a shared text space and align only tokens that occupy identical text spans under the student and teacher tokenizers. To mitigate the problem of excessive generation length and frequent truncation, we introduce a student reference KL loss and mask the advantages of special termination tokens such as </think> and <|im_end|>. This strategy constrains the student from drifting excessively from its initial policy, thereby mitigating the teacher-student distribution mismatch problem and fostering steady length growth. Experiments on both same-family and different-family student models, including Qwen3, Qwen3.5, Intern-S2, GLM-4.7, Gemma-4, show consistent gains in mathematical reasoning, especially natural-language math proving. Notably, Intern-S2-Preview improves by 21.2 points on ProofBench, reaching 55.2 and surpassing Gemini-2.5-Pro. It also improves on science benchmarks such as HLE and HiPhO, suggesting that OPD transfers reasoning capabilities that generalize beyond the mathematical training domain.