How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings
This controlled study demonstrates that while sparse autoencoder features exhibit genuine cross-lingual semantic overlap, their auto-generated natural language labels fail to generalize across different languages, scripts, and rewordings, often reflecting training data representation biases rather than the underlying concepts they are intended to describe.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant, super-smart library of books written in dozens of languages. Inside this library, there are millions of tiny, invisible "switches" (called features) that light up whenever the library reads a specific idea, like "deception," "sports," or "politics."
Because there are too many switches to check by hand, the library uses a robot assistant to look at the top books that make each switch light up and write a little sticky note label for it. For example, if a switch lights up mostly when reading about "Olympic athletes," the robot writes a label: "Olympic Athletes."
This paper asks a simple but critical question: Does that sticky note label tell the whole truth?
Specifically, if the label says "Olympic Athletes," does that switch actually light up for Olympic athletes no matter how the story is told? Does it work if the story is in English? In Russian? In Serbian written in Latin letters? And in Serbian written in Cyrillic letters (a different alphabet)?
Here is what the researchers found, using a clever trick involving the Serbian language.
The "Double-Script" Trick
Serbia is unique because its people write the exact same language in two different scripts: Latin (like English: A, B, C) and Cyrillic (like Russian: А, Б, В). You can turn a Serbian sentence from one script to the other perfectly, like a digital translation, without changing a single word's meaning.
The researchers used this to create a "controlled test." They took 300 sentences and showed them to the AI in four ways:
- English (Latin script)
- Russian (Cyrillic script)
- Serbian (Latin script)
- Serbian (Cyrillic script)
Since the Serbian Latin and Serbian Cyrillic versions are just the same text written differently, the AI should treat them exactly the same. If the "Olympic Athletes" switch is truly about the concept of athletes, it should light up for all four versions.
The Good News: The Switches Do Understand Meaning
First, the researchers checked if the switches themselves actually understood the ideas across languages.
- The Result: Yes! When the AI read the same story in English, Russian, and Serbian, the same group of switches tended to light up.
- The Analogy: Imagine a choir singing a song. Even if they switch from singing in English to singing in Russian, the same group of singers (the features) are still harmonizing on the same melody (the meaning). The switches are genuinely tracking the idea, not just the specific words.
The Bad News: The Sticky Note Labels Lie (Silently)
Here is where it gets tricky. The researchers then looked at the robot-generated labels (the sticky notes). They asked: "If the label says 'Olympic Athletes,' does the switch actually fire when we see that concept in Serbian?"
- The Result: No, not always.
- The labels worked well for English.
- They worked okay for Russian.
- But they failed up to 4 times more often for Serbian, especially when written in the Cyrillic script.
- The Analogy: Imagine a security guard with a badge that says "Detects Fire." He does a great job spotting fires in the main lobby (English). He's pretty good in the back office (Russian). But in the basement (Serbian Cyrillic), he misses the fire completely. The problem isn't that he can't see the fire; the problem is that the badge he's wearing doesn't warn you that he's blind in the basement.
Why Does This Happen?
The paper suggests the labels are biased by what the AI "read" most often during its training.
- The Training Diet: The internet (where AI learns) is mostly English. It has a decent amount of Russian. But it has very little Serbian, and even less Serbian written in Cyrillic.
- The Consequence: The AI is "fluent" in English and "okay" in Russian, but it's a bit "clueless" in Serbian Cyrillic.
- The Label Trap: The robot assistant writes the label based on the examples it sees most often (the English ones). So, it writes a perfect label for English. But because the AI hasn't seen enough Serbian Cyrillic examples, the switch doesn't actually fire for that version of the story. The label is locally accurate (true for English) but globally misleading (false for Serbian).
The "Deep" Problem
The researchers also found that this problem gets worse the deeper you go into the AI's brain (the later layers of the network).
- The Analogy: Think of the AI as a factory assembly line. At the start of the line, the workers (features) are very general and handle all languages well. As the product moves down the line, the workers become specialists. By the end of the line, the specialists are so focused on the "English version" of the product that they stop recognizing the "Serbian version" of the exact same item. The label, however, stays the same, giving you a false sense of security.
The Bottom Line
The paper concludes that auto-generated labels are not guarantees.
- They tell you what a feature does on the most common inputs (like English).
- They do not tell you if that feature works for less common languages or scripts.
- If you rely on a label like "Violent Content" to monitor an AI's safety, you might think you're safe because the label says so. But if the AI encounters that same violent content in a less common language or script, the "safety switch" might not even turn on, and the label gives you no warning that this failure is happening.
In short: The label names the concept correctly, but it hides the fact that the concept might be invisible to the AI in certain forms.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.