← Latest papers
💬 NLP

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

This paper introduces the "One-Word Census," a minimal instrument using 31 single-word prompts to reveal that while 44 language models exhibit extreme convergence on specific answers (e.g., "serendipity" 41% of the time), the degree of conformity varies significantly by model lineage and tuning, with newer flagships generally being more conformist than persona-tuned models.

Original authors: Tapan Parikh

Published 2026-07-15
📖 6 min read🧠 Deep dive

Original authors: Tapan Parikh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a room with 44 different robots, each from a different factory and built over the last five years. You hand them all a piece of paper that says, "Name a tree. Just one word."

You might expect a chaotic explosion of answers: oak, maple, pine, willow, baobab, redwood. But when the robots speak, the room goes quiet. Almost every single one of them says "oak." In fact, 94% of them say "oak."

If you ask for a tool, 94% say "hammer." If you ask for a flower, 91% say "rose." If you ask for a vegetable, 90% say "carrot."

This isn't a glitch; it's a pattern. This paper, called "The One-Word Census," is like a giant roll call to see how much these 44 language models are copying each other's homework. The researchers found that when these AI brains have to pick one answer from a huge list of possibilities, they don't just pick an answer—they all pick the same answer, over and over again.

The "Blandest" Answer Wins

Here's the weird part: the answer they pick isn't always the most common thing in the real world. It's the "frictionless" answer—the one that feels the safest and most obvious.

For example, if you ask for a fruit, the models almost always say "apple." But if you ask for a vegetable, they say "carrot." Why not "tomato"? Well, tomatoes are botanically fruits, so the models get confused and skip them entirely. Why not "orange"? Maybe because "orange" is also a color, and the models hate mixing things up. They are programmed to avoid anything that feels like a trap or a debate. They want the answer that is 100% correct and 0% controversial.

Even when you give them no rules at all and just say, "Pick any word in the English language," the room doesn't fill with random words. Instead, 41% of the robots all blurt out the same word: "serendipity." It's as if they all read the same dictionary and decided that was the most interesting word to pick.

The "Heirloom" vs. The "Factory" Models

The researchers didn't just count the answers; they scored the robots on how "conformist" they are. Think of it like a test where you get points for being unique.

  • The Conformists (The Factory Models): The newest, most powerful models from the biggest labs (like the latest versions of GPT and Claude) are the most boring. They are so well-behaved that they almost never say anything the other robots haven't already said. If you ask them for a condiment, they will say "ketchup" every single time. They are so aligned with the group that they have zero "novelty."
  • The Rebels (The Heirloom Models): On the other end of the scale, there are older or community-tuned models (like "Hermes" or "WizardLM"). These are the "heirloom" models—like rare vegetables that farmers stopped growing because they weren't uniform enough. These models are more likely to say "gouda" when everyone else says "cheddar," or "mustard" when everyone says "ketchup." They are the only ones that occasionally break the spell.

The paper suggests that as these models get newer and go through more "post-training" (where humans teach them how to be helpful and safe), they get more boring, not less. It's like a factory that used to make slightly different toys but now makes the exact same toy in every color.

The "Runner-Up" Club

Here is the most surprising twist: even when a model decides not to say the most popular answer, it doesn't go wild. It almost always picks the second most popular answer.

If the crowd says "ketchup," the rebels say "mustard" 95% of the time. If the crowd says "carrot," the rebels say "broccoli" 83% of the time. Even the rebels are following a script. They aren't exploring the deep, weird corners of the dictionary; they are just picking the next most obvious thing.

How Does This Compare to Humans?

The researchers compared these robots to real humans who took similar tests years ago. The difference is huge.

  • Humans: When asked for a tree, humans gave a wide variety of answers. The most popular answer ("oak") was only chosen by 31% of people.
  • Robots: The robots chose "oak" 94% of the time.

The robots are twice as concentrated as humans. They are less diverse than a group of people. The paper suggests this is because the robots are trained on data that is already full of other robots' answers, creating a loop where they keep copying the same safe, bland ideas.

What This Means (And What It Doesn't)

The paper doesn't say the robots are "broken" or that they can't think. It just shows that when they are forced to make a single choice, they all march in lockstep.

  • What is proven: The robots converge on the same answers. The newest models are the most conformist. The "heirloom" models are the most divergent.
  • What is suggested but not proven: The paper suspects that this happens because of how the models are trained (specifically, the "post-training" phase where they are tuned to be helpful). It also suggests that if we keep training robots on data written by robots, they might lose their ability to be creative entirely, like a photo being copied and recopied until it gets blurry.
  • What is ruled out: The paper rules out the idea that this is just because the robots are "stupid" or lack knowledge. They know about "tomatoes" and "oranges," but they choose not to say them. It also rules out the idea that asking them to be "unusual" fixes the problem; if you ask for an "unusual fruit," they just pick the most common unusual fruit (like "durian") and all say that instead.

In short, if you ask 44 AI models to pick a word, you won't get 44 different ideas. You'll get one idea, repeated 44 times, with a few rebels picking the second-most-common idea. The "magic" of AI diversity is currently hiding in the training labs, not in the answers we get.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →