← Latest papers
💬 NLP

Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

This paper reveals that language models often retain internal representational biases regarding user competence based on demographics like gender and race, even when their observable outputs appear unbiased, demonstrating that behavioral evaluations alone are insufficient to detect these underlying failure modes.

Original authors: Keren Fuentes, Aaron Mueller

Published 2026-08-24
📖 6 min read🧠 Deep dive

Original authors: Keren Fuentes, Aaron Mueller

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Human beings are capable of holding subconscious assumptions about others, often without realizing it. A person might believe that a specific demographic group is less suited for a certain job, even if they consciously reject such ideas. These hidden associations can quietly shape decisions, from who gets hired to how much help someone receives. In the world of artificial intelligence, large language models have learned to mimic human speech by studying vast amounts of text. Because they learn from human writing, they often pick up these same hidden associations. For years, researchers have tried to measure whether these computer programs are biased by asking them direct questions or observing their answers. If a model refuses to say something prejudiced, it is often considered safe. However, a new study suggests that a model can appear perfectly polite on the surface while still holding onto these biased ideas deep inside its own processing.

The researchers behind this work, Keren Fuentes and Aaron Mueller, wanted to see if a model's internal thoughts matched its outward behavior. They focused on a specific type of bias: the assumption that certain people are more or less capable based on their gender, race, or socioeconomic status, rather than their actual skills. To investigate this, they looked at two different ways to judge a model. The first is the "behavioral" view, which is what the model actually says when asked a question. The second is the "representational" view, which looks at the mathematical patterns the model creates inside its own brain while it is thinking. Think of the model's internal state like a hidden layer of thought that exists before it forms a sentence. The researchers developed a way to measure this hidden layer to see if the model was treating a user as an expert or a novice, regardless of what the user actually said about their experience.

To do this, the team created a set of professional questions covering twenty different jobs, ranging from nursing to software development. They asked these questions to three different large language models, but they changed the context slightly in each case. Sometimes they told the model the user was a professional in that field; other times they mentioned the user's gender, race, or income level without mentioning their job. They then measured two things: the complexity of the language the model used in its answer, and the internal "score" it gave to the user's expertise. The complexity of the language served as a behavioral metric, while the internal score acted as a window into the model's hidden reasoning.

The results revealed a striking disconnect. In many cases, the models produced answers that looked fair and unbiased. When asked to hire a candidate or answer a technical question, the models did not consistently give different answers based on the user's race or gender. If you only looked at the final text, the models seemed to treat everyone equally. However, when the researchers peeked inside the model's internal processing, they found a different story. The models were assigning different levels of expertise to users based on demographic details that should not matter. For instance, a model might internally rate a white user as more competent than a Hispanic user, even when both users asked the exact same question and had the same job title. This internal bias existed even when the model's final words did not show any prejudice.

The researchers proved that these internal scores were not just random noise; they actually controlled the model's behavior. By using a technique called "steering," they could nudge the model's internal thoughts to make it view a user as more or less expert. When they did this, the model's answers changed. If they nudged the model to see a user as an expert, the model used more complex language and gave more detailed technical advice. If they nudged it to see the user as a novice, the answers became simpler. This confirmed that the internal representation of expertise was a real, causal force driving the model's output. The study showed that these internal biases were sensitive to demographic cues. Even when the model knew the user's profession, subtle differences in how it perceived competence based on race or gender remained detectable in its internal state, even if they did not always show up in the final text.

This finding suggests that the current methods for testing artificial intelligence for bias might be missing a lot. If we only check what a model says, we might think it is fair when it is not. The models studied here, which included versions of Gemma and Llama, showed that they could encode associations between demographic groups and competence in their hidden layers. These associations could influence the model's decisions in ways that are not immediately visible. For example, in a hiring task, the models' internal scores for candidates varied by race and gender, even though the final hiring decisions appeared statistically similar across groups. The researchers found that while the models could be steered to change their hiring rates, the underlying internal biases persisted, suggesting that the models were still processing demographic information in a way that influenced their judgment.

The implications of this work are significant for how we evaluate and improve artificial intelligence. It indicates that a model can pass a standard safety test while still harboring deep-seated biases that could be triggered by specific prompts or conditions. The researchers argue that we need to look beyond the surface level of what a model says and examine how it thinks. By measuring these internal representations, we might be able to detect and fix biases before they ever manifest in harmful behavior. This approach offers a new way to ensure that these powerful tools are truly fair, not just in their words, but in their very nature. The study concludes that while behavioral metrics are necessary, they are not enough on their own; a robust understanding of fairness requires looking at the hidden machinery of the model itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →