← Latest papers
💬 NLP

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

The paper introduces ADAGE, a language-agnostic pipeline that creates translation-free analogical reasoning benchmarks in Arabic, Amharic, and Japanese, revealing a significant performance gap where models proficient in English struggle with culturally grounded reasoning in these native languages.

Original authors: Ahmed Haj Ahmed, Alvin Grissom II

Published 2026-07-28
📖 5 min read🧠 Deep dive

Original authors: Ahmed Haj Ahmed, Alvin Grissom II

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to understand human jokes. You might think the best way is to take a book of famous jokes written in English, translate them into Spanish, Japanese, or Arabic, and ask the robot to explain them. But here's the catch: a joke often relies on the specific culture it was born in. If you translate a joke about a specific American holiday into a language where that holiday doesn't exist, the joke loses its magic. The robot might guess the answer because it sounds funny, not because it actually understands the meaning. This is the problem scientists face when testing "multilingual" AI. They often just translate English tests into other languages, which creates a fake illusion that the AI is smart in those languages.

To truly test if an AI can think like a human, we need to see if it can understand the deep, unspoken rules of a culture—like knowing that "don't count your chickens before they hatch" means you shouldn't be too confident until something actually happens. This kind of thinking is called analogical reasoning. It's the ability to take an old idea and apply it to a brand-new situation. If an AI can only do this in English, but fails when the same logic is presented in a different culture, it hasn't really learned to think; it's just memorized English patterns. This paper asks a simple but scary question: Are our AI models actually smart in other languages, or are they just really good at translating?

The researchers introduce a new tool called ADAGE (Analogical Difficulty-by-design Assessment for Grounded Evaluation). Think of ADAGE as a master chef who doesn't just translate recipes from a French cookbook; instead, they go into local markets in Arabic, Amharic, and Japanese to find the freshest, most authentic local ingredients (proverbs) and cook up brand-new dishes (scenarios) that test if the AI can taste the difference.

Here is how they cooked up their experiment:

  1. The Ingredients: They gathered thousands of proverbs (short, wise sayings like "He who does not know the falcon will roast it") from native speakers in Arabic, Amharic, and Japanese.
  2. The Recipe: Instead of just asking the AI to explain the proverb, they used a clever trick. They grouped proverbs by their "flavor" or theme (like "patience" or "caution"). Then, they asked an AI to write a modern story for each proverb.
  3. The Trap: To make the test hard, they created a multiple-choice question. They gave the AI a story and four proverbs to choose from. One was the perfect match. The other three were "distractors":
    • One was from the same theme group (a "fake-out" that sounded right but wasn't quite right).
    • One was a different theme but sounded similar (a "near-miss").
    • One was completely random (the "easy" wrong answer).
  4. The Taste Test: They fed these tests to 14 different AI models, ranging from tiny ones to massive ones, and watched how they performed.

What They Found

The results were a bit of a shock, like finding out a student who aced the English test failed the math test in a different language.

  • The "Translation Trap" is Real: When the AI models were tested on English proverbs, they did very well, scoring around 80% or higher. But when the researchers tested them on the native Arabic, Amharic, and Japanese benchmarks, the scores plummeted. The best models dropped by 12 to 52 percentage points. For example, a model that got 83% right in English might only get 70% right in Arabic, and a mere 41% in Amharic.
  • The "Low-Resource" Reality: The Amharic test was the hardest. Even the biggest, smartest models barely scored better than random guessing (which would be 25% for a four-choice question). The AI models seemed to get stuck on the language itself, unable to even understand the words, let alone the deep meaning.
  • Size Isn't Everything: The researchers expected that bigger models (with more "brain power") would do better. While bigger models did improve on Arabic and Japanese, the improvement was small. On Amharic, making the model bigger didn't help at all. It suggests that just adding more data or size doesn't fix the problem if the AI doesn't have a deep cultural foundation to stand on.
  • The Distractor Trick Worked: The clever "distractor" strategy worked perfectly. The AI models made the most mistakes on the "near-miss" answers (the ones that sounded similar but were wrong), proving they were actually trying to reason and not just guessing randomly.

The Big Takeaway

The paper concludes that the current way we test AI—by simply translating English tests into other languages—is lying to us. It makes AI look smarter in other languages than it really is. The AI models are great at English cultural reasoning, but they struggle to apply that same logic to the rich, complex cultures of Arabic, Amharic, and Japanese.

The authors suggest that until we build tests that are native to each culture, using local proverbs and stories, we won't know if our AI is truly multilingual or just a very good translator. They have released their new "kitchen" (the ADAGE pipeline) and their "recipes" (the benchmarks) so other scientists can cook up their own tests for any language in the world, ensuring that future AI can truly understand us, no matter where we are from.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →