Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages
This paper introduces Onramp-Sequence Cross-Distillation (OSCD), a post-training algorithm that combines an integrated translator agentic loop with joint-embedding semantic alignment to overcome cross-lingual collapse and catastrophic forgetting, significantly enhancing native mathematical reasoning capabilities in low-resource Southeast Asian languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where your favorite video game character could only speak English, no matter what language you spoke to them. If you asked a question in Spanish, they'd understand the words, but when they started to explain their answer, they'd suddenly switch back to English for the hard thinking parts. This isn't just a glitch; it's a major hurdle in the world of Artificial Intelligence, specifically for "Large Language Models" (LLMs). These are the super-smart computer brains that can write stories, solve math problems, and chat with us. To get really good at solving hard puzzles, they use a trick called "Chain-of-Thought," which is like whispering their step-by-step thinking process to themselves before giving the final answer. The problem is, for many languages around the world—especially those with fewer digital resources like those in Southeast Asia—these AI brains tend to "collapse" and whisper their thoughts in English, even when you asked them to think in Thai, Vietnamese, or Indonesian. This makes the AI feel less like a helpful local friend and more like a distant, English-only tutor, which is a big deal for education and fairness.
Enter a new study by researchers Sean Gip Lim, William Chandra Tjhi, and Hai Leong Chieu, who decided to build a bridge for these AI brains. They created a clever training method called OSCD (Onramp-Sequence Cross-Distillation). Think of it like a bilingual coach who doesn't just translate the final answer, but actually rewrites the AI's internal "thinking notes" in the local language while keeping the correct solution intact. Usually, when you try to teach an AI a new language, it gets confused and forgets how to be smart in its original language (a problem called "catastrophic forgetting"). But this new method uses a special "translator agent" to swap the thinking steps and a "semantic glue" to make sure the AI's brain still understands the logic, just in a different tongue.
The researchers tested this on some of the toughest math puzzles available, like the AIME25 and HMMT25 competitions, which are like the Olympic Games for math problems. They took powerful AI models that were great at English math but terrible at thinking in Southeast Asian languages and gave them a short, intense training session using only 70,000 examples (which is tiny compared to the millions usually needed). The results were impressive: the models didn't just get better at answering the questions; they actually started thinking in the local languages. In fact, the models improved their ability to solve these math problems while staying in the target language by up to 3.2 times. Even more exciting, they managed to do this without losing their original English smarts. The study suggests that by using this "translator agent" and the "semantic glue," we can finally give AI the ability to reason naturally in languages that have been left behind, making advanced technology accessible to everyone, not just English speakers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.