← Latest papers
💬 NLP

IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages

This paper introduces IndicIFEval, a benchmark comprising 800 human-verified examples across 14 Indic languages to evaluate LLMs' instruction-following capabilities, revealing that while models adhere well to formatting constraints, they significantly lag behind English performance in lexical and cross-lingual tasks.

Original authors: Thanmay Jayakumar, Mohammed Safi Ur Rahman Khan, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan

Published 2026-02-26
📖 5 min read🧠 Deep dive

Original authors: Thanmay Jayakumar, Mohammed Safi Ur Rahman Khan, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, multilingual robot assistant. You can ask it to write a story, summarize a news article, or solve a math problem. But there's a catch: you also want to give it very specific rules, like "Write exactly three sentences," "Don't use the word 'cat'," or "Start your answer with the word 'Hello'."

This paper, IndicIFEval, is like a giant, rigorous "driver's license test" for these robot assistants, but specifically for 14 different Indian languages (like Hindi, Bengali, Tamil, and Sanskrit).

Here is the breakdown of what they did, using some simple analogies:

1. The Problem: The "English-Only" Bias

Imagine a driving school that only tests you on English roads. You might be a perfect driver in English, but if you move to a country where they drive on the left, speak a different language, and have different traffic signs, you might crash.

Currently, most AI tests are built in English. The researchers found that while these AI models are great at following rules in English, they often get confused, forget instructions, or make up their own rules when asked to do the same thing in Indian languages. There was no good way to measure how bad this was, so they built a new test.

2. The Solution: Two Types of Tests

To make the test fair and thorough, they created two different "exam rooms" (datasets):

  • Room A: The "Translation" Exam (INDICIFEVAL-TRANS)

    • The Analogy: Imagine taking a famous English exam and translating it word-for-word into Hindi or Tamil.
    • The Goal: This checks if the AI can handle the structure of the instruction. However, the researchers were careful. They didn't just use Google Translate. They had human experts tweak the questions so they made sense culturally. For example, changing "The President of the USA" to "The Prime Minister of India" so the AI doesn't get confused by the context.
    • The Catch: Sometimes, a direct translation feels "stiff" or unnatural, like wearing a suit that was tailored for someone else's body.
  • Room B: The "Native" Exam (INDICIFEVAL-GROUND)

    • The Analogy: Instead of translating an English exam, imagine writing brand-new questions from scratch, based on local culture, news, and stories that naturally happen in India.
    • The Goal: This tests if the AI can follow rules when the instructions feel like they were written by a local person, not a translator. It's like asking the AI to write a poem about a local festival with specific rhyming rules, rather than translating a poem about Christmas.

3. The Test Subjects: The AI Models

They put over 20 different AI models through these tests. Some were "Open-Weight" (like open-source software anyone can download and tweak) and some were "Proprietary" (like the secret, high-end models from big tech companies).

They tested models of all sizes, from tiny ones (like a pocket calculator) to massive ones (like a supercomputer).

4. The Results: What Happened?

The results were a mix of good news and bad news:

  • The "Formatting" Superpower: The AI models were surprisingly good at following "shape" rules. If you asked them to "write in a JSON format" or "use bullet points," they usually got it right. They are like students who are great at following the layout of a form but struggle with the content.
  • The "Word" Struggle: The models struggled hard with "word" rules. If you asked them to "include the word 'elephant' exactly three times," they often forgot, used it twice, or used a synonym instead. It's like a student who knows the essay topic but keeps forgetting to write their name on the paper.
  • The "English Gap": Even the best models performed significantly worse in Indian languages than in English. It's like a musician who can play a concerto perfectly in a major key (English) but fumbles the notes when asked to play in a complex, local scale (Indic languages).
  • The "Thinking" Boost: They found that when they told the AI to "think step-by-step" before answering (a mode called "Reasoning"), it got much better at following the rules. It's like giving the student a few extra minutes to read the instructions twice before starting the test.
  • The Winners: The Gemma family of models (from Google) seemed to have the best "multilingual muscle memory," performing closest to English levels. The Proprietary models (the expensive, closed ones) generally did better than the open ones, but the open ones were catching up fast.

5. Why Does This Matter?

This paper is a wake-up call. It shows that while AI is getting smarter, it still has a "cultural blind spot." If we want AI to be truly helpful for billions of people who speak Indian languages, we can't just translate English tests. We need to build tests that understand the rhythm, grammar, and culture of those languages.

In a nutshell: The researchers built a specialized gym for AI to practice following strict rules in Indian languages. They found that while the AI is getting stronger, it still trips over its own feet when the language gets complex or the instructions get specific. But, with the right training (and "thinking" time), it can learn to walk the tightrope.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →