← Latest papers
💬 NLP

UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop

This paper introduces UrduBench, a standardized Urdu reasoning benchmark created through a contextually ensembled translation framework with human-in-the-loop validation, which reveals significant challenges in multi-step and symbolic reasoning for large language models while establishing a scalable methodology for evaluating low-resource languages.

Original authors: Muhammad Ali Shafique, Areej Mehboob, Layba Fiaz, Muhammad Usman Qadeer, Hamza Farooq

Published 2026-01-30
📖 4 min read☕ Coffee break read

Original authors: Muhammad Ali Shafique, Areej Mehboob, Layba Fiaz, Muhammad Usman Qadeer, Hamza Farooq

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant student who can solve complex math problems and answer tricky questions in English. Now, imagine you want to test if this student can do the same things in Urdu, a language with fewer digital resources. The problem is, if you just ask a machine to translate the English test questions into Urdu, the translation might be clunky, lose its meaning, or even change the rules of the game. The student might fail not because they aren't smart, but because the test itself is broken.

This paper, titled "UrduBench," is about building a fair, high-quality test for Large Language Models (LLMs) in Urdu. Here is how they did it and what they found, explained simply:

1. The Problem: The "Broken Translation" Trap

Think of translating a logic puzzle like translating a recipe. If you translate the ingredients list separately from the cooking instructions, the final dish might be inedible. Similarly, existing methods often translated Urdu test questions word-by-word without looking at the whole picture. This caused "context fragmentation," where the question and the answer choices didn't fit together logically anymore.

2. The Solution: The "Translation Team" Approach

The authors didn't just use one translator. They built a contextually ensembled framework, which is like hiring a team of four different expert translators (IndicTrans2, NLLB, Qwen, and Gemini) to work on the same question at the same time.

  • The Team: They translated the entire question and all its answer choices together as one block, ensuring the story stayed intact.
  • The Editor: They used a super-smart AI (GPT-5.1) to act as an editor, comparing the four translations and stitching together the best parts of each one.
  • The Human Check: Finally, real human experts who are fluent in both English and Urdu reviewed the results. They acted as the "quality control inspectors," picking the most natural and accurate version.

This process created UrduBench, a collection of four famous reasoning tests (Math, Common Sense, Science, and Logic) translated perfectly into Urdu.

3. The Experiment: The "Urdu Olympics"

Once they had the fair test, they invited various AI models to take it. They tested models of different sizes (from tiny 1-billion-parameter models to massive 12-billion-parameter ones) and different types (some trained specifically for reasoning, others just for following instructions).

They asked the models to solve problems in three ways:

  • Direct: Just give the answer.
  • Chain-of-Thought (CoT): "Show your work" step-by-step.
  • Few-Shot: "Here are some examples, now you try."

4. The Results: What the Models Learned

The paper found several interesting things, which can be summed up with these metaphors:

  • Math is Harder than Common Sense: The models struggled much more with multi-step math problems (like the MGSM and MATH-500 tests) than with common sense questions. It's like the models can easily tell you "a cat has fur," but they trip over complex algebra in Urdu.
  • Bigger Isn't Always Better: Usually, a bigger engine means a faster car. But in Urdu, a slightly smaller model (Gemma-3-4B) actually drove better than some larger ones. This suggests that how the model was trained (specifically its exposure to Urdu) matters more than just raw size.
  • The "Reasoning" Specialist vs. The "Generalist": Models specifically designed to be "reasoning experts" did great on math but didn't always beat the general "instruction-following" models on common sense questions. Being a specialist in math didn't automatically make them a master of all Urdu tasks.
  • The Language Confusion Test: This was a crucial finding. The researchers checked if the models answered in pure Urdu or if they accidentally slipped back into English (code-switching). They found that models that stayed strictly in Urdu performed much better. If a model got confused and mixed languages, its reasoning skills dropped. It's like a chef who starts speaking a different language while cooking; the dish might get ruined.

5. The Conclusion

The paper concludes that to build smart AI for low-resource languages like Urdu, you can't just translate English tests and hope for the best. You need a human-in-the-loop process to ensure the translation preserves the logic.

They also found that for an AI to reason well in Urdu, it needs to be linguistically consistent. If the AI loses its way and starts mixing languages, its brain fog sets in, and it fails the test.

In short: The authors built a high-quality, human-verified "Urdu Olympics" for AI. They discovered that while AI is getting better, it still struggles with complex math in Urdu, and the key to success is keeping the language pure and the translation context-aware. They have made their test and data available for others to use.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →