← Latest papers
💬 NLP

MAPS: A Multilingual Benchmark for Agent Performance and Security

MAPS is a new multilingual benchmark suite that evaluates the performance and security of agentic AI systems across eleven languages by translating tasks from four major benchmarks (GAIA, SWE-Bench, MATH, and Agent Security), revealing that transitioning from English to other languages often leads to degraded reliability and increased security risks.

Original authors: Omer Hofman, Jonathan Brokman, Oren Rachmil, Shamik Bose, Vikas Pahuja, Toshiya Shimizu, Trisha Starostina, Kelly Marchisio, Seraphina Goldfarb-Tarrant, Roman Vainshtein

Published 2026-02-11
📖 3 min read☕ Coffee break read

Original authors: Omer Hofman, Jonathan Brokman, Oren Rachmil, Shamik Bose, Vikas Pahuja, Toshiya Shimizu, Trisha Starostina, Kelly Marchisio, Seraphina Goldfarb-Tarrant, Roman Vainshtein

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a highly skilled personal assistant—someone who can book your flights, fix your computer code, solve complex math problems, and keep your digital life secure. Now, imagine this assistant is a genius in English, but the moment you speak to them in Spanish, Hindi, or Japanese, they suddenly become clumsy, forgetful, and—worst of all—dangerously unpredictable.

That is the problem this research paper, MAPS, is trying to solve.

The Problem: The "Language Glass Ceiling"

Current AI "agents" (AI that doesn't just talk, but actually does things) are mostly trained and tested in English. Because the underlying "brains" of these agents (Large Language Models) are often better at English, the agents inherit a hidden weakness: they struggle when the instructions come in other languages.

It’s like hiring a world-class chef who can follow a recipe perfectly in English, but if you give them the same recipe in French, they might accidentally swap salt for sugar or forget to turn on the oven. In the digital world, a "forgotten oven" could mean a bank transaction goes wrong or a security door is left unlocked.

The Solution: MAPS (The Multilingual Stress Test)

The researchers created MAPS, which stands for a "Multilingual Benchmark." Think of MAPS as a Global Proficiency Exam for AI agents.

Instead of just asking the AI simple questions, they put the agents through four intense "job simulations":

  1. The Real-World Assistant (GAIA): Can the agent browse the web and find information to solve a real task?
  2. The Software Engineer (SWE-Bench): Can the agent find and fix bugs in computer code?
  3. The Mathematician (MATH): Can the agent solve high-level math problems?
  4. The Security Guard (ASB): Can the agent resist being "tricked" or "hacked" by bad actors?

They didn't just do this in English. They translated these tests into 11 different languages (like Arabic, Chinese, German, and Hindi) to see how the agents would hold up globally.

The "Aha!" Moment: What They Found

The researchers discovered something critical: The more "human" and "wordy" a task is, the more the AI trips up.

  • The "Math & Code" Exception: When the tasks were mostly numbers or computer code (which look similar in almost every language), the agents were fine. It’s like a musician playing a sheet of music—the notes are the same regardless of what language you speak.
  • The "Natural Language" Trap: When the tasks required understanding subtle human instructions (like "Find me a flight that isn't too loud"), the agents' performance plummeted.
  • The Security Risk: Most alarmingly, they found that the agents became less secure in other languages. A "jailbreak" (a trick to make the AI do something bad) that wouldn't work in English might work perfectly in another language. It’s like a security guard who knows how to spot a thief in a suit, but lets a thief walk right past them if they are wearing a different style of clothing.

Why This Matters to You

As we move toward a world where AI agents handle our money, our schedules, and our data, we can't afford for them to have a "language bias."

This paper is a wake-up call to AI developers. It says: "If you want your AI to be a reliable global citizen, you can't just teach it to be a genius in English; you have to make sure it stays a genius in every language its users speak."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →