← Latest papers
📈 economics

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

This study demonstrates that while large language models can generate vast numbers of survey respondents, their resulting "silicon surrogates" produce highly stylized and distorted representations of human cultural tastes that systematically overestimate liking, fail to capture complex relational structures, and misalign with key social dimensions like age, class, gender, and race.

Original authors: Xiangyu Ma, Mengmi Zhang, Shannon Ang, Minne Chen

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Xiangyu Ma, Mengmi Zhang, Shannon Ang, Minne Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand what music people actually like. You could ask 10,000 real humans, or you could ask a super-smart computer program to pretend to be 10,000 humans. This paper asks: If you ask the computer, will it tell you the truth?

The researchers took a real survey about music tastes (the SPPA) and asked three different AI models (from OpenAI, Anthropic, and DeepSeek) to generate "fake" answers for every single person in that survey. They then compared the AI's answers to the real humans' answers to see how close the imitation was.

Here is what they found, explained simply:

1. The AI is a "Hyper-Enthusiast" (The Positive Bias)

Imagine a person who, when asked if they like a specific type of music, says "Yes!" to almost everything. That is the AI.

  • The Finding: The AI respondents said they liked music genres far more often than real people do. If real people liked 24% of the genres on average, the AI said they liked about 46%.
  • The Metaphor: Real humans are picky eaters; they might like pizza but hate broccoli. The AI is a "super-eater" who claims to love everything on the menu, including the weird, obscure dishes that most people actually dislike. The researchers call this "hyper-omnivorousness."
  • The Twist: You might think the AI is bad at understanding "minority" groups because it was trained mostly on data from wealthy, educated, white people (the "WEIRD" bias). But the study found the opposite: The AI was actually worse at simulating wealthy, educated, white people than it was at simulating others. It just liked everything, regardless of who it was pretending to be.

2. The AI Misses the "Vibe" (The Relationality Problem)

Human taste isn't just a list of "Yes" and "No." It's a web of connections. For example, if you love Jazz, you are likely to also love Blues. If you love Country, you might also like Bluegrass. These connections create a "vibe" or a structure.

  • The Finding: The AI failed to understand these connections. It didn't know that certain genres "go together."
  • The Metaphor: Imagine a human music lover's taste is like a well-organized bookshelf where genres are grouped logically. The AI's taste is like a pile of books thrown on the floor. It might have the same number of books as the human, but the way they are arranged makes no sense.
  • The Result: The AI broke the links between genres that usually go together (like Blues and Classic Rock) and invented fake links between genres that never hang out together (like Hip-Hop and Bluegrass). It looked like a music lover on the surface, but the internal logic was completely scrambled.

3. The AI Gets the "Social Map" Wrong

In the real world, what music you like often depends on who you are (your age, your income, your gender, your race).

  • Age: The AI thinks almost all music is for young people. It made genres like "Classic Rock" (which older people usually love) seem like they were for teenagers.
  • Class (Money): The AI invented a strict "rich vs. poor" divide that doesn't really exist anymore. It decided that "High Brow" music (like Classical and Opera) is strictly for the rich, and "Low Brow" music is for the poor, ignoring the fact that real people today mix these tastes much more freely.
  • Gender: The AI turned subtle differences into loud stereotypes. If real women slightly prefer Broadway musicals, the AI made it look like a massive, defining trait. It also erased real trends, like the fact that women actually like Country music quite a bit.
  • Race: This was the biggest failure. The AI invented fake racial associations. It decided that some genres were "for white people" and others were "for Black people" in ways that were exaggerated or completely made up. For example, it made Hip-Hop look more racially charged than it is in reality, and it even invented a racial link for Electronic music where none existed before.

The Bottom Line

The paper concludes that while AI can mimic the surface of human opinion (it can say "I like music"), it fails to capture the depth and structure of human taste.

Think of the AI as a caricature artist. If you ask a caricature artist to draw a person, they might get the nose and the eyes right, but they will exaggerate the features so much that the drawing looks funny and distorted. It looks like a human, but it's not faithful.

If researchers use these AI "surrogates" to study culture, they will get a distorted picture of the world: a world where everyone likes everything, where music genres don't connect logically, and where social groups are defined by exaggerated, fake stereotypes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →