← Latest papers
💬 NLP

Multi-lingual Functional Evaluation for Large Language Models

This paper introduces multi-lingual functional benchmarks (CL-GSM Symbolic and CL-IFEval) across five languages to demonstrate that static multi-lingual evaluations often overestimate model performance and robustness compared to practical, task-based assessments.

Original authors: Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian

Published 2026-03-13
📖 4 min read☕ Coffee break read

Original authors: Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new employee for a global company. You want to know if they can actually do the job in different languages, not just if they can pass a written test in English.

This paper is like a report card for that hiring process, but instead of a person, the "employee" is an AI (Large Language Model).

Here is the story of what the researchers found, explained simply:

1. The Old Way: The "Multiple Choice" Trap

For a long time, we tested AI models using Static Benchmarks. Think of these like standardized multiple-choice tests (like the SATs or a trivia quiz).

  • How it works: The AI reads a question in French or Spanish and picks the right answer from a list.
  • The Problem: The researchers found that these tests are like cheating. An AI might memorize the answers to the questions in the test bank during its training. It gets a high score, but it doesn't actually understand how to use the language in the real world. It's like a student who memorized the answer key but can't hold a conversation.

2. The New Way: The "Functional" Test

The authors created two new tests called CL-GSM Symbolic and CL-IFEval. Think of these as practical driving tests or on-the-job simulations.

  • How it works: Instead of asking "What is 2+2?", they give the AI a template with moving parts.
    • Example: "Sally bought {number} red apples and {number} green apples. How many total?"
    • The AI has to ignore the colors (the distraction) and focus on the math (the function).
  • The Twist: They translated these tests into five languages: French, Spanish, Hindi, Arabic, and Yoruba (a language spoken in Nigeria).

3. The Big Surprise: The "Hype vs. Reality" Gap

The researchers ran the tests and found a massive disconnect between the "Multiple Choice" scores and the "Driving Test" scores.

  • The Illusion: On the old tests, the AI looked like a genius in all languages.
  • The Reality: When put to work with the new functional tests, the AI's performance crashed in many languages.
    • In English, French, and Spanish, the AI's math skills dropped by about 17% to 24% when switching from the old test to the new one.
    • It's like a student who gets an 'A' on the written exam but fails the practical lab because they can't actually mix the chemicals.

4. The "Resource" Divide

The paper highlights a cruel reality about AI: It loves rich languages and struggles with poor ones.

  • High-Resource Languages (English, French, Spanish): The AI is decent here. It has seen a lot of data on the internet.
  • Low-Resource Languages (Yoruba): The AI is practically blind here. In the new functional tests, the AI's performance on Yoruba was drastically lower (sometimes below 15% accuracy).
  • The Analogy: Imagine a library. The English section is a massive, well-lit library with millions of books. The Yoruba section is a single, dusty pamphlet in a dark corner. The AI is a student who has read the English library cover-to-cover but has never seen the pamphlet.

5. The "Brittle" AI

The most important finding is about robustness.

  • Some AI models are like glass: They look perfect until you tap them with a slightly different question, and then they shatter.
  • The researchers found that an AI might be great at following instructions in English but completely fail at the exact same instruction in Arabic or Hindi.
  • The "Start/End" Failure: In one specific test, the AI failed to follow a simple rule ("Start your sentence with a capital letter") in Yoruba, even though it could do it perfectly in English. It's like a robot that can dance to jazz but freezes if you ask it to dance to salsa.

Summary: What Does This Mean for Us?

This paper is a wake-up call.

  • Don't trust the hype: Just because an AI says it's "multilingual" and scores high on standard tests doesn't mean it can actually do things in those languages.
  • The "English Bubble": AI is currently very good at English but is still very fragile and unreliable in many other languages, especially those with fewer digital resources (like Yoruba).
  • The Fix: We need to stop testing AI with static quizzes and start testing it with real-world functional tasks to see if it can actually handle the job.

In a nutshell: The AI is a great student who memorized the textbook, but when you ask it to solve a real-life problem in a different language, it often forgets how to think. We need better tests to find out who is actually ready for the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →