On-Policy Delta Distillation for Multilingual Math Reasoning
This paper demonstrates that On-Policy Delta Distillation (OPD), which leverages the probability gap between a post-trained teacher and its base model, significantly enhances multilingual mathematical reasoning in Korean and Japanese while narrowing performance gaps, though it underscores the necessity of multilingual data to prevent English-centric response shifts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers are like brilliant, multilingual students who have read every book in the library but still struggle to solve a tricky math problem. To help them get smarter, scientists use a technique called "training," which is a bit like a coach guiding a student through practice problems. For a long time, the best way to train these AI students was through Reinforcement Learning (RL), a method where the computer gets a single grade at the very end of its answer—like getting an "A" or "F" on a whole essay without knowing which specific sentences were good or bad. However, a newer, more efficient method called On-Policy Distillation (OPD) has emerged. Instead of waiting for a final grade, OPD acts like a master teacher who reads the student's answer word-by-word, whispering corrections and suggestions as the student writes. This "token-level" feedback is much more precise. But here's the catch: most of this research has only been done in English. We didn't know if this "whispering teacher" technique worked just as well for students speaking Korean, Japanese, or other languages, or if it could help bridge the gap between how well the AI speaks English versus how well it speaks other tongues.
This paper dives into that exact question, testing a supercharged version of the technique called On-Policy Delta Distillation (OPD²) on mathematical reasoning in English, Korean, and Japanese. The researchers found that this new method is a game-changer. Just like a master chef who can teach a recipe to a student in any language, OPD² helped the AI models solve math problems significantly better in all three languages, with particularly huge leaps in Korean and Japanese. The study suggests that this method is even better than the original version because it focuses on the "delta"—the specific difference between what the teacher knows and what the student's base model already knows—effectively isolating the new reasoning skills being taught.
However, the paper also uncovered a fascinating, slightly tricky side effect. When the researchers tried to teach the AI using only English math problems, the model surprisingly got better at solving Korean and Japanese math problems too. It was as if the student learned the logic of the math in English and could apply it to other languages. But there was a twist: while the answers were correct, the model often started writing its explanations in English, even when the question was asked in Korean or Japanese. This suggests that while the "reasoning muscle" can be trained in one language and transferred to another, keeping the "voice" of the answer in the target language requires practicing with data in that specific language. Ultimately, the paper shows that OPD² is a powerful tool for making AI smarter across languages, but to truly master a language, you still need to speak it during practice.
The Core Discovery: A Better Whispering Teacher
The researchers set out to see if OPD² (the advanced version of the whispering teacher) could outperform the original OPD when dealing with multiple languages. They used a family of AI models called Qwen3 (specifically the 1.7B and 8B versions) and a massive teacher model to guide them. The results were clear and consistent: OPD² consistently beat the original OPD across the board.
Think of it like this: If the original OPD was a good tutor, OPD² was a genius tutor who knew exactly what the student was missing. In the experiments, this difference was huge for non-English languages. For the smaller model (Qwen3-1.7B), OPD² improved the Korean math score from 40.9 to 51.9, and the Japanese score from 37.2 to 52.0. The original OPD helped, but OPD² pushed those numbers even higher. For the larger, smarter model (Qwen3-8B), the improvements were just as strong, with Korean jumping from 57.0 to 64.4 and Japanese from 57.1 to 65.0.
The paper explicitly argues that this isn't just about copying the teacher's style. By using the "delta signal" (the gap between the teacher and the teacher's own base model), OPD² successfully transferred the reasoning capabilities to the student, rather than just mimicking the teacher's general output. This suggests that the method is robust and works well regardless of the language or the size of the model.
Bridging the Language Gap
One of the biggest goals of this research was to see if these training methods could narrow the performance gap between English and other languages. Often, AI models are much better at English than at Korean or Japanese. The paper found that multilingual OPD² generally helped narrow this gap.
In a detailed look at the "English–Korean accuracy gap," the researchers saw that while the base model favored English by a wide margin (sometimes by over 13 points on certain tests), the OPD training reduced this disparity. For instance, on the Global-MGSM benchmark, the gap shrank from 13.4 points down to 8.6 points with standard OPD, and further down to 9.3 points with OPD². On the HRM8K benchmark, the gap dropped from 12.1 to 9.2 points.
The paper suggests that this happens because the multilingual training benefits the non-English languages more than the English language, effectively lifting the weaker performers without dragging down the strong ones. However, the authors are careful to note that this effect wasn't perfectly uniform across every single test; on some specific benchmarks like M-MMLU, the gap actually widened slightly. So, while the trend is positive, it's not a magic wand that fixes every single metric instantly.
The "English-Only" Surprise and the Language Trap
Perhaps the most surprising finding in the paper is what happened when they trained the models using only English data. You might think that if you only teach a student English math, they would fail at Korean math. But the paper found that English-only OPD² actually improved the Korean and Japanese scores significantly.
In non-thinking mode, the English-only training boosted the Korean average from 40.9 to 52.6 and the Japanese average from 37.2 to 53.4. These numbers were almost as good as the multilingual training scores (which were 51.9 for Korean and 52.0 for Japanese). This suggests that the logic of mathematical reasoning is so universal that it can be learned in English and then applied to Korean or Japanese problems, even if the model never saw a single Korean math problem during training.
However, there is a major catch. The paper explicitly rules out the idea that this means the model has truly mastered the language. When the researchers checked what language the model actually spoke in its answers, the results were stark.
- Multilingual Training: When trained on mixed languages, the model answered in Korean 90.5% of the time (in non-thinking mode) and in Japanese 90.9% of the time.
- English-Only Training: When trained only on English, the model's "Korean" answers dropped to 48.3% (meaning it answered in English half the time), and its "Japanese" answers plummeted to 29.6%.
Even in "thinking mode" (where the model does a hidden reasoning step before answering), the trend held. The English-only model would solve the math correctly but often write the final answer in English, ignoring the fact that the question was asked in Korean or Japanese. The paper concludes that while English-only training can transfer the ability to reason, it fails to preserve the habit of responding in the target language. To get the model to speak the language you want, you must feed it data in that language.
Summary of Findings
The paper demonstrates that OPD² is a superior method for teaching AI mathematical reasoning across languages, consistently outperforming the original OPD. It successfully narrows the performance gap between English and languages like Korean and Japanese, suggesting that multilingual post-training is a viable path to more balanced AI. However, the study also draws a clear line between "reasoning ability" and "language generation." While reasoning skills can be learned in one language and applied to another, the model will naturally drift toward English unless it is explicitly trained with data in the target language. The authors suggest that future research should focus on evaluating both the accuracy of the answer and the language of the response to truly understand multilingual AI capabilities.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.