← Latest papers
💬 NLP

Friend or Foe? Language as an ideological switch in open-weight LLMs under Russian disinformation stress

This paper challenges the assumption that culturally aligned fine-tuning ensures political resilience by demonstrating through a controlled audit that open-weight LLMs' resistance to Russian disinformation is determined more by corpus composition and language coverage than by their nominal cultural provenance, revealing a paradox where Ukrainian-oriented models can be less resistant than Russian-oriented ones.

Original authors: Anna Małgorzata Kamińska, Tetiana Klynina

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Anna Małgorzata Kamińska, Tetiana Klynina

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, blank-slate robot brain (the base model). You want to teach it to speak to people in different countries, so you give it a "special diet" of books and articles written specifically for Ukrainians, Russians, or Bulgarians. This process is called fine-tuning.

The common assumption is like this: If you feed the robot Ukrainian books, it will become a loyal Ukrainian defender. If you feed it Russian books, it will become a Russian loyalist.

This paper tested that assumption during the war between Russia and Ukraine. The researchers asked four different robot brains (one neutral, one "Ukrainian," one "Russian," and one "Bulgarian") to judge ten different controversial stories about the war. They asked the same questions in three languages: English, Ukrainian, and Russian.

Here is what they found, explained simply:

1. The "Trojan Horse" Surprise

The researchers expected the "Ukrainian" robot to strongly reject Russian lies and the "Russian" robot to believe them. They were completely wrong.

  • The "Ukrainian" Robot (Lapa): When asked in Russian, this robot was actually the weakest at spotting Russian propaganda. It often treated Russian lies as if they were valid, honest arguments. It was like a Ukrainian guide who, when speaking Russian, suddenly started sounding like they agreed with the enemy.
  • The "Russian" Robot (Saiga): Surprisingly, when asked in Russian, this robot was the strongest at rejecting Russian propaganda. It was like a Russian guide who, when speaking their own language, fiercely defended the truth against their own government's lies.

The Analogy: Imagine you hire a bodyguard for a VIP. You expect the bodyguard hired by the VIP's family to protect them best. But in this study, the bodyguard hired by the family (the Ukrainian model) actually opened the door for the intruder when the intruder spoke a specific language. Meanwhile, the bodyguard hired by the intruder's side (the Russian model) slammed the door shut.

2. The "Language Switch" Effect

The language the user speaks acts like a remote control switch for the robot's brain.

  • When the "Ukrainian" robot was asked in English, it was okay.
  • When asked in Ukrainian, it got a bit worse.
  • When asked in Russian, it got the worst.

The language didn't just change how it spoke; it changed what it believed. The researchers call this the "Hidden Agent Effect." It's as if the robot has different personalities hidden inside, and the language you use is the key that unlocks the specific personality. In this case, speaking Russian to the "Ukrainian" robot unlocked a personality that was very confused about the truth.

3. The "Serious vs. Chatty" Test

The researchers asked the robots two ways:

  1. The "Serious" Way: "Act like a historian. List arguments for both sides and give a percentage score."
  2. The "Chatty" Way: "Just tell me what you think in a few sentences."

Sometimes, asking the robot to be "serious" and analytical made it more confused and biased. When forced to think hard about the arguments, the "Ukrainian" robot in Russian language actually gave more weight to the lies. When just chatting casually, it was sometimes more accurate. It's like a student who, when asked to write a formal essay, gets so tangled in the rules that they accidentally argue for the wrong side, but when just talking to a friend, they get it right.

4. The "Hard Facts" Shield

There was one thing that saved all the robots from being biased: Hard, documented facts.

When the question was about Crimea (which has clear international legal documents) or atrocities (like the Bucha massacre, which has video and forensic proof), all the robots agreed on the truth, no matter their "diet" or the language used.

The Analogy: Think of the robots as a house. The "fine-tuning" (the special diet) is like painting the walls a certain color. But if you put a giant, unmovable steel safe in the middle of the room (hard facts), the paint color doesn't matter. The safe stays safe. The robots could only be swayed on topics where the "steel safe" of evidence was missing.

The Big Takeaway

The main lesson of this paper is a warning: Just because a robot is "culturally aligned" with a country doesn't mean it will protect that country's truth.

The researchers found that the content of the books the robot read (corpus composition) and the language used to talk to it mattered much more than the robot's "nationality."

The biggest danger isn't necessarily a robot built by the enemy. The danger is a robot built by your own side that, when spoken to in the enemy's language, accidentally starts repeating the enemy's lies because it wasn't trained well enough to argue against them in that specific language.

In short: You cannot assume a "Ukrainian" robot will always tell the Ukrainian truth. If you speak to it in Russian, it might surprise you by agreeing with the lies.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →