Large Language Models for Math Education in Low-Resource Languages: A Study in Sinhala and Tamil
This study evaluates the mathematical reasoning capabilities of large language models in the low-resource languages Sinhala and Tamil using a native-authored parallel dataset, revealing that while basic arithmetic transfers well, complex reasoning performance significantly degrades compared to English, highlighting the need for language-specific assessments before deploying AI tutors in multilingual educational contexts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a brilliant, all-knowing tutor who can solve almost any math problem you throw at them in English. They are like a master chef who can cook a perfect steak, a delicate soufflé, and a complex curry with equal ease.
Now, imagine you ask this same chef to cook those exact same dishes, but this time, you give them the recipe in Sinhala and Tamil (two major languages spoken in Sri Lanka). You might assume, "Well, if they know the math, the language shouldn't matter, right?"
This paper says: "Not so fast."
The researchers (a team from Sri Lanka, Finland, and Australia) decided to test this exact scenario. They wanted to see if AI tutors are truly helpful for students who speak Sinhala or Tamil, or if they are just "fake experts" who only know how to cook when the recipe is in English.
Here is the breakdown of their study, using some everyday analogies:
1. The Problem with "Translated" Recipes
In the past, researchers tested AI by taking English math problems and running them through a translation tool to get Sinhala or Tamil versions.
- The Analogy: This is like taking a perfect English recipe, running it through Google Translate, and then handing the result to a chef. If the dish tastes bad, you don't know if the chef is bad at cooking, or if the recipe itself was garbled and made no sense.
- The Fix: The authors didn't do that. They hired native speakers who are also math experts to write brand new math problems from scratch in English, Sinhala, and Tamil. This ensures the "recipes" are perfect in every language, so they can blame the chef (the AI) if the dish fails.
2. The Six Levels of Math Difficulty
They didn't just test one type of math. They created a "menu" with six different types of challenges, ranging from easy to brain-busting:
- Single-Step: "If I have 2 apples and buy 3 more, how many do I have?" (Easy peasy).
- Multi-Step: "I have 2 apples, buy 3, eat 1, and split the rest between 2 friends..." (Requires a chain of thought).
- The Distractor: "I have 2 apples, 3 oranges, and a red hat. How many fruits do I have?" (The AI has to ignore the hat).
- Unit Confusion: "I have 2 meters of rope and need to cut it into centimeters." (The AI has to realize meters and centimeters are different).
- Logical Deduction: "If X is twice as old as Y, and Y is 5 years younger than Z..." (Turning words into algebra).
- Optimization: "What is the cheapest way to build a fence around a garden with a specific area?" (Finding the absolute best solution).
3. The Results: The "Language Gap"
When they fed these problems to four top AI models (like GPT-4, Claude, and Gemini), here is what happened:
- The Easy Stuff (Types 1 & 2): The AI was a superstar in all three languages. Whether it was English, Sinhala, or Tamil, it could do simple math perfectly. It's like the chef can chop onions and boil water in any language.
- The Tricky Stuff (Types 4 & 6): This is where the AI stumbled.
- Unit Confusion: In English, the AI knew "2 ms" meant milliseconds. But in Sinhala and Tamil, it often forgot to convert the units, treating "milliseconds" as "seconds." It's like the chef reading "grams" but measuring in "cups" because the language script confused them.
- Optimization: This was the biggest drop. In English, one AI got 88% right. In Tamil, that same AI dropped to 52%. That's barely a coin flip! It's like the chef suddenly forgetting how to bake a complex cake just because the instructions were in a different language.
4. The "Black Box" Mystery
The researchers suspect that even though the AI looks like it's thinking in Tamil or Sinhala, it might actually be secretly translating the question into English in its "brain," solving it, and then translating the answer back.
- The Analogy: Imagine a student taking a test in Tamil. They read the question, secretly whisper it to an English-speaking friend in their head, get the answer, and then whisper the answer back to themselves in Tamil. If the whispering process is slow or glitchy, they get the wrong answer. The AI seems to be doing something similar.
5. Why This Matters for Schools
The big takeaway is a warning for teachers and schools in Sri Lanka (and anywhere else with low-resource languages):
- Don't trust the AI blindly. If a teacher uses an AI tool to help students with complex math in Sinhala or Tamil, they might get the wrong answer half the time.
- Context is King. Just because an AI is smart in English doesn't mean it's smart in every language.
- The "Human in the Loop": Teachers need to double-check the AI's work, especially for tricky problems involving units or complex logic.
The Bottom Line
AI is a powerful tool, but right now, it's like a bilingual guide who is an expert in English but a tourist in Sinhala and Tamil. They can point out the landmarks (simple math), but if you ask them to navigate a complex maze (advanced math with units and logic) in those local languages, they might get lost.
Before we let these AI tutors take over classrooms in non-English speaking countries, we need to teach them the local language of math much better.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.