← Latest papers
💬 NLP

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

This paper demonstrates that in large language models, the properties of demographic identity being readable, faithfully structured, and causally used are three dissociable phenomena, revealing that while specific attention heads can encode survey-like group differences with high fidelity, this encoding does not necessarily translate into the model's actual causal usage for generating responses.

Original authors: Fathin Difa Robbani

Published 2026-08-20
📖 6 min read🧠 Deep dive

Original authors: Fathin Difa Robbani

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Large language models are increasingly being asked to act as stand-ins for human beings in social science research. Instead of spending months and millions of dollars interviewing thousands of real people to understand how different groups feel about politics, religion, or economics, researchers can simply ask a computer program to simulate those responses. The hope is that these digital surrogates can capture the subtle differences between a young Democrat and an older Republican, or between a Hindu and a Christian, just as a real survey would. But a growing body of evidence suggests these simulations are failing. The computer's answers tend to be too uniform, too agreeable, and too similar to one another, regardless of the demographic label attached to the prompt. The central mystery has been whether this failure happens because the model simply does not know the differences between these groups, or because it knows the differences but chooses not to use them when it speaks.

To solve this, a researcher set out to look inside the "brain" of a large language model to see where demographic information actually lives. The investigation focused on a specific model, Mistral-7B, and asked three distinct questions. First, is the information about who is answering readable from the model's internal signals? Second, is that information arranged in a way that mirrors the real world, where the distance between two groups in the computer's mind matches the distance between them in reality? And third, does the model actually use this internal map when it generates an answer? The study treated these three questions as separate properties, realizing that a model could possess a faithful internal map without ever consulting it to form an opinion.

The researchers began by mapping the model's internal landscape. They created prompts for 169 different intersectional groups, such as young Asian Democrats or middle-aged Hindu Republicans, and examined the model's internal state at the very end of its processing. They compared the model's internal arrangement of these groups against real survey data from the Pew Research Center, which provided the ground truth of how these groups actually differ in their opinions. The results revealed that the standard way of reading the model's mind—looking at the final output layer—was misleadingly weak. It suggested the model had a blurry, indistinct view of demographic differences. However, when the researchers looked deeper, at the specific components that make up the model's processing, they found a much clearer picture.

The information was not lost; it was just hidden in specific, narrow channels. The study discovered that the model's internal geometry is dominated by individual attention heads, which are specialized sub-processors within the model. In five out of six types of demographic attributes tested, the best single attention head provided a much more accurate map of group differences than the model's entire final output layer. One specific attention head, located in the 11th layer of the model, stood out as remarkably faithful. It held a representation of all six demographic types—age, education, income, religion, political party, and political ideology—that closely matched the real-world structure of how these groups differ. This single component carried a map where the relationships between groups were arranged with a fidelity that reached roughly 70% of the theoretical limit of what could be measured, even after accounting for the fact that the survey data itself contains some noise.

Yet, finding a faithful map inside the model did not mean the model was using it. The researchers then tested whether this accurate internal arrangement actually influenced the model's answers. They performed a causal experiment where they swapped the demographic identity in the prompt and watched how the model's predicted opinions shifted. The result was stark: changing the entire identity of the person in the prompt moved the model's predicted answers by less than 2% of their total error. The model produced slightly different, but equally wrong, answers for different groups. Even more surprisingly, the researchers found that the location with the most faithful map was not the location that drove the answers. The clearest causal pathway, where changing the internal state actually moved the model's prediction toward the truth, was found in a demographic type that had one of the weakest internal maps. Conversely, the most faithfully arranged type showed no detectable causal effect when the researchers tried to intervene on it.

This disconnect suggests that the failure of these models to simulate populations is not simply a lack of knowledge. The model does hold a detailed, accurate map of how different groups differ, particularly regarding political and economic factors, while its understanding of race and religion remains weak and unstable. However, the model does not consult this map when it speaks. The internal geometry is faithfully arranged but causally disconnected from the final output. The study also found that simply reading this internal map directly, using a mathematical probe on that single faithful attention head, produced results that were 21% to 31% closer to real survey truth than the model's own generated answers. However, even this superior read-out could not predict the specific order of opinions for individual questions any better than the model's own flawed answers.

The implications of these findings are significant for anyone trying to use artificial intelligence to understand human society. The research demonstrates that the ability to simulate a population cannot be assumed just because a model can be prompted to act like a specific person. The model possesses the knowledge, but it is not using it in the way researchers hope. Furthermore, the study warns that demographic attributes are not interchangeable; while the model has a strong grasp of political and economic identities, its understanding of racial and religious identities is fragile and prone to breaking under slight changes in how a question is asked. The work concludes that treating the model's internal knowledge and its external behavior as the same thing is what keeps the debate over whether these models can simulate populations unresolved. The model has the map, but it is not following the route.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →