用于多语言数学推理的同策略增量蒸馏
On-Policy Delta Distillation for Multilingual Math Reasoning
August 6, 2026
作者: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
cs.AI
摘要
在线策略蒸馏(OPD)正成为强化学习在LLM后训练中的一种有前景的替代方案,然而其在多语言场景下的有效性仍研究不足。我们研究了OPD及其进阶变体——在线策略增量蒸馏(OPD²),用于英语、韩语和日语的数学推理任务。OPD²通过利用后训练教师模型与其基础模型之间的概率差距作为学习信号,改进了OPD。基于Qwen3的实验表明,OPD²始终优于原始OPD,尤其在韩语和日语上提升显著,并且总体上缩小了英语与韩语之间的性能差距。我们进一步发现,仅使用英语数据的OPD也能提升韩语和日语的性能,但往往会使回复偏向英语,这凸显了多语言数据对于保持目标语言回复的重要性。
English
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.