ChatPaper.aiChatPaper

多言語数学推論のためのオン方策デルタ蒸留

On-Policy Delta Distillation for Multilingual Math Reasoning

August 6, 2026
著者: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
cs.AI

要旨

オンポリシー蒸留(OPD)は、LLMのポストトレーニングにおける強化学習に代わる有望な手法として注目されているが、多言語環境での有効性はまだ十分に検討されていない。本研究では、英語・韓国語・日本語における数学的推論に対して、OPDとその発展版であるオンポリシーデルタ蒸留(OPD^2)を検証する。OPD^2は、ポストトレーニング済みの教師モデルとそのベースモデル間の確率ギャップを学習信号として用いることで、OPDを改善する。Qwen3を用いた実験では、OPD^2が元のOPDを一貫して上回り、特に韓国語と日本語で顕著な改善が見られ、英語-韓国語間の性能ギャップも概ね縮小することが示された。さらに、英語のみのOPDでも韓国語と日本語の性能が向上する可能性があるが、応答が英語に偏ることが多く、対象言語の応答を維持するには多言語データの重要性が強調される。
English
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.