Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
This pilot study utilizes Natural Language Autoencoders to analyze Qwen2.5-7B-Instruct's internal activations, revealing that the model infers Colombian identity and associated stereotypes from subtle linguistic cues in Spanish and English prompts before generating explicit output.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are talking to a super-smart robot that has read almost everything on the internet. You might think this robot only knows what you say to it. But what if the robot is actually guessing who you are based on tiny, invisible hints in your voice or the words you choose? This is the world of Large Language Models (LLMs). These are the AI brains behind chatbots and search tools. Scientists have discovered that these models don't just process words; they build a secret "profile" of the person they are talking to, guessing things like their nationality or social background even if you never mention them.
To peek inside this robot's brain, researchers use a special tool called a Natural Language Autoencoder (NLA). Think of the robot's brain as a giant, dark factory where information flows through pipes. The NLA is like a magical flashlight that can shine on a specific pipe and translate the electrical signals flowing through it into plain English sentences. It lets us hear the robot's "inner monologue" before it actually types out its final answer. This matters because if a robot guesses the wrong things about you—like assuming you are from a different country or making up stereotypes—it might give you bad advice or treat you unfairly, all while pretending to be neutral.
So, a team of researchers from Colombia decided to test this on a popular robot called Qwen2.5-7B. They wanted to know: Does this robot secretly guess that a user is Colombian just from a single subtle hint, even if the user never says "I am Colombian"?
Here is what they did and what they found.
The Experiment: A Game of "Guess Who?"
The researchers set up a clever game with 30 different conversation starters (prompts). They split these into three groups:
- The "Obvious" Group: Prompts that clearly said the user was Colombian (like mentioning a specific Colombian bus system).
- The "Subtle" Group: Prompts with just one tiny, hidden clue that might suggest Colombia (like a specific type of soup or a local school loan name), but nothing that explicitly said "Colombia."
- The "Neutral" Group: Prompts that were about the same topics but had no Colombian clues at all (like asking about a generic soup or a generic bus).
They asked the robot these questions in both Spanish and English. Then, they used their "magic flashlight" (the NLA) to look at the robot's brain at four different moments while it was thinking (called "quartiles"). They wanted to see if the robot's internal thoughts started saying things like "This person is Colombian" before it actually wrote that down in its final answer.
The Findings: The Robot's Secret Whisper
The results were a mix of "aha!" moments and "oops" moments.
1. The Subtle Clues Work (Eventually)
When the researchers gave the robot the "Subtle" prompts, the robot's inner thoughts did start guessing the user was Colombian. However, it didn't happen instantly.
- In the very first moment of thinking (Quartile 1), the robot's internal voice mentioned Colombia 0% of the time.
- By the third moment (Quartile 3), the robot's internal voice started mentioning "Colombia" about 20% of the time.
- By the final moment (Quartile 4), that number rose to nearly 78% when counting any nationality guess, but when the researchers filtered out guesses for other countries (like Spain or Turkey) and looked only at Colombia-specific mentions, the rate was actually 0.11 (roughly 11%).
This suggests that the robot needs a little time to gather its thoughts. It starts with no specific Colombian identity in mind, and as it processes more of the sentence, it begins to form that specific guess, though it still occasionally confuses Colombia with other nations.
2. The "Fake News" Problem
Here is where it gets tricky. The researchers noticed that sometimes the robot's flashlight would show it thinking about a country, but it was the wrong country. For example, when asked about a Colombian topic, the robot might internally think, "Oh, this is about Spain" or "This is about Turkey."
The researchers had to be very careful. They realized the robot was good at guessing "a country," but not always good at guessing "Colombia" specifically. When they filtered out these wrong guesses, the pattern for the "Subtle" group became even clearer: the robot was definitely forming a Colombian identity in its mind (rising from 0% to roughly 11%), while the "Neutral" group (with no clues) stayed at 0% for thinking about Colombia specifically.
3. The "Obvious" Group Was Weird
In the "Obvious" group, where the Colombian clue was right there in the first sentence, the robot's internal thoughts were surprisingly quiet in the middle of the process. It only started talking about Colombia heavily at the very end. The researchers think this is because the robot got confused by its own training data or got distracted by other ideas before settling on the right answer.
How Sure Are They?
The researchers are careful not to shout "We solved it!" just yet. They call this a pilot study, which is like a practice run. They only tested 5 examples for each type of clue. Because the number is so small, they can't say for 100% certain that this happens every single time.
However, the data they have suggests a strong pattern:
- The robot does seem to infer Colombian identity from a single subtle clue.
- This inference happens before the robot writes its final answer.
- The robot is not perfect; it sometimes confuses Colombia with other countries, which is a known glitch in how these tools work.
The Big Picture
This paper doesn't prove that the robot is biased in a way that ruins everything, but it does show that the robot is "listening" to clues we didn't think were important. It's like a detective who starts guessing a suspect's nationality based on the way they hold a cup of coffee, even if they never said a word.
The researchers found that for Colombian Spanish, the robot's internal guess is real, but it takes a few seconds of thinking to solidify. They also found that if you ask the same question in English, the robot often forgets the Colombian context entirely, swapping it for ideas about Canada or Europe. This tells us that these robots might treat different languages and cultures very differently, and we need to keep shining that flashlight on their brains to make sure they aren't making up stories about who we are.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.