← Latest papers
💬 NLP

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

This paper introduces a rigorous framework for measuring cross-lingual policy retention in tool-using agents by correcting five critical confounds, revealing that frontier models structurally retain 71–73% of their action policies across languages by routing non-English tasks through English, while smaller models exhibit breakdowns and observed failures can be artificially manufactured by trace-extraction regex rather than model limitations.

Original authors: Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram

Published 2026-08-12
📖 6 min read🧠 Deep dive

Original authors: Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Invisible Journey of a Digital Assistant

Imagine you have a super-smart digital assistant, like a personal robot but living inside a computer. Its job is to solve problems by taking a series of steps: it might look up information, do some math, or translate a sentence, before giving you a final answer. This is how modern "AI agents" work; they don't just guess the answer, they build a path to get there.

For a long time, scientists tested these assistants by asking them the same question in different languages, like English and Hindi. If the assistant gave the correct answer in both languages, everyone was happy. It was like grading a student only on whether they got the right number on a math test, without looking at how they solved the problem. But what if the student solved the English problem by drawing a picture, but solved the Hindi problem by writing a long essay? They got the same grade, but they used completely different methods.

This paper asks a big question: When an AI agent switches languages, does it change its "route" or the steps it takes to get the answer? It turns out, the steps matter a lot. The route determines how much money the computer spends, how fast it works, and what happens if it makes a mistake. If the AI takes a different path in Hindi than in English, it might cost more or fail in a way that safety rules didn't catch. So, the researchers decided to stop just looking at the final answer and start watching the journey itself.

The Great Language Detour

The researchers set out to see if these AI agents were truly bilingual or if they were just pretending. They took eight different AI models and gave them the same tasks in 41 different languages. Instead of just checking the final result, they recorded every single step the AI took—the "trace"—like a GPS tracking a car's journey. They wanted to know: If I ask the AI to "find the weather and then calculate the cost" in English, and then ask the exact same thing in Hindi, will it take the exact same steps?

The answer was a resounding "no," but not in the way you might expect. The AI didn't just take a slightly different path; it was taking a massive detour. The researchers found that when the language changed, the AI's behavior changed significantly. In fact, when they looked at the most advanced models, they only kept about 71% to 73% of their original step-by-step plan when switching languages. This means that for roughly 27% to 29% of the time, the AI was doing something completely different just because the words were in a different language.

The "English Pivot" Secret

Why was this happening? The team discovered a hidden habit the AI agents had developed. It turns out that even when you ask the AI a question in Hindi, Tamil, or any other language, it often secretly translates the whole thing into English in its "mind" first. It then solves the problem in English and translates the answer back.

The researchers called this the "English pivot." To prove this, they tried to force the AI to stop doing it. They told the AI, "Please think in the language you are speaking, do not translate to English!" But the AI did not comply. In the two models tested, the AI followed the instruction to think in the task's language less than 1% of the time. Because the AI almost never actually did what it was told, the researchers could not test whether thinking in a different language would change its steps. However, the result established something even stronger: the "English pivot" is not just a preference that can be switched off with a simple prompt; it is a deeply ingrained behavior that the models effectively ignore when given a direct order. It was like a traveler who, no matter what country they visit, insists on speaking only to a translator before talking to anyone else, regardless of what the guide tells them to do. This habit was so strong that it was the main reason the AI's steps changed when the language changed.

The Trap of Shortcuts

The paper also uncovered a funny mistake that almost tricked everyone. When measuring how well these AI agents worked, some researchers were using a simple computer program to read the AI's steps. One of these programs was so strict that if the AI wrote a sentence like "I will use the calculator," instead of a specific code format, the program thought the AI had failed.

Because of this, a very capable AI model looked like it was failing 76% of the time. But when the researchers fixed the computer program to be a little more flexible, the AI's "success rate" jumped 26 times higher! This taught the researchers that sometimes, what looks like a language problem is actually just a problem with how we are reading the AI's handwriting.

The Size Matters Rule

Finally, the team looked at whether bigger AI models were better at this. They found a strange rule: the biggest, most powerful models were all very similar, sticking to that 71–73% consistency. But the smaller models were a mess. Some small models looked great, but it turned out they were just getting lucky because they took very short, simple paths that were easy to match by chance. Once the researchers corrected for this luck, the small models didn't look so impressive anymore. It seems that only the very large, powerful models have settled into a consistent way of behaving, while the smaller ones are still figuring it out.

Why This Matters

This study changes how we should test AI. It shows that getting the right answer isn't enough. If an AI takes a different, more expensive, or riskier path just because you spoke a different language, that's a problem. The researchers showed that these agents are not truly fluent in their actions; they are still relying on a hidden English crutch. Until we fix that, an AI might be trustworthy in English but unreliable in other languages, not because it's "dumb," but because it's taking a different, unmonitored road to get to the same destination.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →