ChatPaper.aiChatPaper

온-폴리시 델타 증류를 통한 다국어 수학 추론

On-Policy Delta Distillation for Multilingual Math Reasoning

August 6, 2026
저자: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
cs.AI

초록

온-폴리시 증류(On-Policy Distillation, OPD)는 LLM 사후 훈련에서 강화학습의 유망한 대안으로 부상하고 있으나, 다국어 환경에서의 효용성은 여전히 충분히 탐구되지 않았다. 본 연구는 영어, 한국어, 일본어 수학적 추론에서 OPD와 그 고급 변형인 온-폴리시 델타 증류(OPD²)를 분석한다. OPD²는 사후 훈련된 교사 모델과 기본 모델 간의 확률 격차를 학습 신호로 활용하여 OPD를 개선한다. Qwen3 실험 결과, OPD²는 원래의 OPD보다 일관되게 우수한 성능을 보였으며, 특히 한국어와 일본어에서 두드러진 향상을 나타냈고 전반적으로 영어-한국어 성능 격차를 축소하였다. 또한 영어 전용 OPD도 한국어와 일본어의 성능을 향상시킬 수 있지만, 응답이 영어로 전환되는 경향이 자주 나타나 목표 언어 응답을 보존하는 데 다국어 데이터의 중요성을 강조하였다.
English
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.