← Latest papers
💬 NLP

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

This paper introduces UrduMMLU, a comprehensive benchmark of over 26,000 Urdu-language multiple-choice questions derived from native educational sources to evaluate the performance of 30 large language models, revealing significant gaps in their understanding of regionally grounded content and humanities subjects compared to STEM.

Original authors: Ahmer Tabassum, Sarfraz Ahmad, Hasan Iqbal, Owais Aijaz, Momina Ahsan, Preslav Nakov

Published 2026-06-08
📖 5 min read🧠 Deep dive

Original authors: Ahmer Tabassum, Sarfraz Ahmad, Hasan Iqbal, Owais Aijaz, Momina Ahsan, Preslav Nakov

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to test how smart a group of students is. For years, you've only been able to give them tests written in English. You know they can read English, but you have no idea if they actually understand the history, literature, and science of their own culture, or if they just memorized English translations of those concepts.

This paper introduces URDUMMLU, a massive new "exam" designed specifically for the Urdu language, spoken by over 230 million people. Here is the breakdown of what they did and what they found, using some simple analogies.

1. The Problem: The "Translated Menu" Issue

Until now, if researchers wanted to test AI models on Urdu, they mostly took English tests and translated them.

  • The Analogy: Imagine trying to judge a chef's ability to cook authentic Italian food by giving them a menu translated from French. They might get the words right, but they might miss the cultural nuances, the local ingredients, or the specific way a dish is traditionally prepared.
  • The Reality: Existing Urdu benchmarks were like these translated menus. They didn't capture the unique flavor of Urdu education, literature, or local history (like Pakistan Studies or Islamic Studies).

2. The Solution: Building a Native "Urdu Exam Hall"

The authors built URDUMMLU, a benchmark with 26,431 multiple-choice questions.

  • Where did the questions come from? Instead of translating English questions, they went straight to the source: actual Pakistani school textbooks, past exam papers (SSC/HSSC), and local Urdu question banks.
  • The "Double-Check" System: Since some of these old exam papers didn't have answer keys, the team hired 17 native Urdu speakers (mostly with university degrees) to act as "proctors."
    • The Rule: Two proctors had to look at every question and agree on the answer. If they disagreed, or if the question was broken, the question was thrown out. This ensured the "gold standard" answers were 100% reliable.
  • The Scope: The exam covers 26 subjects, ranging from Math and Physics (STEM) to Urdu Literature, Islamic Studies, and Pakistan Studies.

3. The Test: Putting 30 AI Models in the Exam Hall

The researchers took 30 different AI models (both famous "closed" ones like Google's Gemini and open-source ones) and gave them this Urdu exam. They tested them in two ways:

  1. English Instructions: "Here is a question in Urdu, pick the answer." (The instruction is in English).
  2. Urdu Instructions: "یہاں ایک سوال ہے، جواب منتخب کریں۔" (The instruction is in Urdu).

4. The Results: The "STEM vs. Culture" Gap

The results were revealing, like a report card that showed a student was great at math but terrible at reading poetry.

  • The Top Performer: The model Gemini-3.5-Flash was the clear winner, scoring over 90%. It was the only model to break the 85% barrier.
  • The Open-Source Gap: The best open-source model (DeepSeek-V4-Flash) scored around 82%. While good, it trailed the leader by about 8 points.
  • The "Humanities" Crash: This is the most important finding.
    • The Analogy: Imagine a student who can solve complex physics equations (STEM) but fails miserably when asked to analyze a poem or explain a religious text (Humanities).
    • The Reality: Almost every model performed much better on Science and Math questions. When the questions shifted to Urdu Literature, Urdu Language, and Islamic Studies, many models dropped 25 to 40 percentage points.
    • Example: A model might get 97% on Physics but only 67% on Urdu Literature. This suggests the models know the facts of science (which are universal) but lack the deep cultural and linguistic understanding required for local subjects.

5. Other Surprising Findings

  • Language of Instructions Didn't Matter Much: Whether the AI was told to answer in English or Urdu, the score barely changed. The problem wasn't the instruction language; it was the lack of knowledge in the model's "brain" about Urdu culture.
  • Few-Shot Learning (The "Cheat Sheet"): The researchers tried giving the AI a few example questions before the test (like a cheat sheet). It helped a little bit (raising scores by 1–3 points), but it wasn't enough to fix the big gaps in cultural knowledge.
  • The "Urdu-Specific" Models: Interestingly, models specifically trained just for Urdu (like Qalb and Alif) performed quite poorly (around 35%), often worse than general models. This suggests that simply training on Urdu text isn't enough; the models need to be trained on educational and reasoning material, not just chat or news.

Summary

URDUMMLU is like a new, strict "driver's license test" for AI in Urdu. It proves that while AI is getting very good at answering science questions in Urdu, it is still struggling to understand the soul of the language—its literature, history, and local culture. The current "best" AI (Gemini) is a strong driver, but the open-source models are still learning the rules of the road, especially when it comes to the cultural nuances of the Urdu-speaking world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →