基於同策略 Delta 蒸餾的多語言數學推理
On-Policy Delta Distillation for Multilingual Math Reasoning
August 6, 2026
作者: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
cs.AI
摘要
同策略蒸餾(On-Policy Distillation, OPD)正逐步成為大型語言模型後訓練中強化學習的可行替代方案,然而其在多語言情境下的有效性仍未獲得充分探討。我們研究了 OPD 及其進階變體——同策略Delta蒸餾(On-Policy Delta Distillation, OPD²)——在英語、韓語與日語數學推理任務上的表現。OPD² 透過利用後訓練教師模型與其基礎模型之間的機率差距作為學習訊號,進而改善了 OPD。基於 Qwen3 的實驗顯示,OPD² 的表現持續優於原始 OPD,在韓語與日語上尤其有顯著進步,且普遍縮小了英語與韓語之間的效能差距。我們進一步發現,僅使用英語資料的 OPD 也能提升韓語與日語的效能,但往往會使回應偏向英語,這凸顯了多語言資料對於維持目標語言回應的重要性。
English
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.