Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles
This paper proposes a methodological framework that treats large language models as implicit sociological tools to reconstruct aggregate voting behavior from sociodemographic profiles, demonstrating their ability to accurately replicate real-world election outcomes and political structures while highlighting their utility as exploratory instruments for computational social science.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the study of how people vote, political scientists have long relied on a simple truth: who you are often predicts how you vote. A person's age, education, income, and where they live create a demographic profile that correlates strongly with their political choices. For decades, researchers have mapped these connections using surveys, asking thousands of people about their lives and their ballots to find the patterns that shape elections. But a new question has emerged in the digital age: could a computer program, trained on the vast ocean of human writing found on the internet, already know these patterns without ever asking a single person a question? Large language models, the powerful artificial intelligence systems behind modern chatbots, are built by reading billions of documents. In doing so, they absorb the statistical regularities of human culture, including the unspoken links between a person's background and their political leanings. This raises a fascinating possibility: could these models act as a compressed mirror of society, holding a hidden map of how demographics translate into votes, simply because they have read so much about it?
A team of researchers from the Czech Republic set out to test this idea using the 2021 Czech parliamentary election as their laboratory. They did not ask the artificial intelligence to predict the winner of the election or to explain why a specific individual chose a certain party. Instead, they treated the models as a kind of sociological instrument, a way to reconstruct the collective behavior of a nation from the ground up. The researchers took a real survey containing nearly 4,000 Czech citizens, each with a detailed profile of their age, gender, education, income, and even their interest in politics. They fed these descriptions, one by one, into four different versions of a large language model. For each person, the model was asked a simple question: given this specific life story, who would you vote for? The model did not just pick one party; it calculated the probability of voting for every major party, along with the chance that the person would not vote at all.
The researchers then combined these thousands of individual guesses into a single national picture. They used a method that treated every probability as a partial vote, allowing the final result to emerge from the sum of many small uncertainties rather than a single forced choice. When they compared this simulated election to the actual official results, the match was surprisingly close. The models reproduced the final vote shares with an average error of only about 2.5 percentage points, a level of accuracy comparable to high-quality pre-election polls conducted by professional agencies. More importantly, the simulation did not just get the numbers right; it captured the hidden structure of Czech politics. The models correctly identified that certain parties naturally cluster together, forming two distinct political blocs that mirror the real-world alliances and rivalries seen in the country. They also recreated the specific ways voters flow between parties, showing that the artificial intelligence understood, without being told, that voters for one coalition often shift to another, while voters for opposing camps rarely cross the divide.
The study revealed that the size of the model mattered. The larger, more capable versions of the artificial intelligence were better at distinguishing between different types of voters, capturing the subtle differences in how a young, educated person might vote compared to an older, less educated one. The smaller models tended to guess the average for everyone, missing the nuances of the demographic landscape. Yet even the smaller models managed to recover the broad outlines of the political map. The researchers also checked what the models knew before they were given any specific person's details. They found that the models already possessed a strong, built-in knowledge of the election results, likely because the outcome was discussed so frequently in the text they were trained on. However, the true value of the experiment was not in this pre-existing knowledge, but in the ability to break that knowledge down. By feeding in individual profiles, the researchers could see how the models distributed that knowledge across different regions and social groups, effectively turning a static fact into a dynamic, structural map of society.
Crucially, the authors emphasize that this is not a tool for predicting the future or for replacing human surveys. The models do not understand why people vote the way they do; they have no causal understanding of human decision-making. They are simply reflecting the patterns they have seen in the text they were trained on. If the internet is full of articles linking education to a specific party, the model learns that link. The value of this work lies in its ability to act as a diagnostic tool, allowing scientists to interrogate what a model has absorbed from human culture. It shows that these systems have internalized the sociological regularities of a non-English speaking, multi-party democracy, even though they were trained mostly on English data. The researchers conclude that while these models are not perfect mirrors of reality, they are powerful enough to reconstruct the skeleton of a society's political behavior, offering a new way to explore the hidden connections between who we are and how we vote.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.