← Latest papers
💬 NLP

Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement

This study demonstrates that while supplying one to three high-signal demographic attributes in prompts can improve alignment between LLMs and human annotations, over-specification with full attribute sets degrades performance, revealing that successful demographic prompting depends on the joint consideration of signal learnability, directional coherence, and task context rather than simply increasing the volume of attributes.

Original authors: Mahammed Kamruzzaman, Shrabon Kumar Das, Gene Louis Kim

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Mahammed Kamruzzaman, Shrabon Kumar Das, Gene Louis Kim

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot friend (a Large Language Model, or LLM) who is really good at reading stories and guessing how people feel about them. Sometimes, the robot guesses perfectly, but other times, it misses the mark because it doesn't know who is reading the story.

Researchers from the University of South Florida decided to test a popular idea: "What if we tell the robot exactly who the reader is?" They tried giving the robot a "persona" by listing the reader's age, race, gender, religion, and more. They wanted to see if this helped the robot agree better with actual human readers.

Here is what they found, using a mix of five different robot brains and five different types of tricky tasks (like spotting toxic comments, guessing if a text is polite, or figuring out the emotion).

The "Goldilocks" Rule: Less is More

The biggest surprise? More information didn't make the robot smarter; it made it confused.

Think of it like giving directions to a friend. If you say, "Go to the park," they might get lost. If you say, "Go to the park, turn left at the red house, then right at the blue car," they might get there. But if you keep adding details—"and don't forget the dog, and the time is 3 PM, and the weather is rainy, and the park has a fountain, and..."—your friend might just freeze up and go the wrong way.

The researchers found that for most robots, the sweet spot was giving them one to three specific details about the reader.

  • The Sweet Spot: When they gave just a few clues (like "a person who is LGBTQ+ and Black"), the robot's guesses matched human readers much better.
  • The Overload: When they gave the robot every single detail available (the "full attribute set"), the robot's performance actually got worse. It was like trying to listen to five people talking at once; the robot couldn't figure out who to listen to.

The "Signal vs. Noise" Mystery

You might think, "If a group of people really disagrees about something, the robot should definitely know about it!" But the paper suggests that's not always true.

Imagine the robot is a detective trying to solve a crime based on clues.

  • The Magnitude Trap: Just because a group of people argues loudly about a topic (high "magnitude" of difference) doesn't mean the robot can use that argument to solve the case. The researchers found that knowing how much humans disagreed didn't help predict if the robot would get better.
  • The Coherence Key: The real secret was directional coherence. This is a fancy way of asking: "Do the different people in this group agree on what the clues mean?"
    • Good Signal: If all the people in a group (say, "LGBTQ+ individuals") agree that the word "threat" means something specific, the robot can learn that.
    • Bad Signal: If half the group thinks "threat" means danger, but the other half thinks it means something else, the robot gets confused. Even if the group is "loud" about their opinions, if they are pulling in opposite directions, telling the robot "act like this group" actually makes it worse.

The "Neuron" Detective Work

To understand why this happened, the researchers peeked inside the robot's brain (a process called "neuron probing"). They looked at which tiny parts of the robot's brain lit up when they gave it a persona.

They found a strange paradox with one specific robot, DeepSeek.

  • The High-Volume Paradox: When DeepSeek was given a persona, it lit up more neurons than any other robot. You'd think, "Wow, it's working so hard!" But actually, it was the worst at following the instructions.
  • The Lesson: Just because a robot's brain is buzzing with activity doesn't mean it's thinking clearly. Sometimes, all that noise is just the robot getting confused by the extra instructions, not actually understanding the persona.

What This Means for You

The paper doesn't say "Demographic prompting is broken." Instead, it suggests that it's a tool that needs to be used carefully.

  • Don't dump everything on the robot: If you want the robot to sound like a specific type of person, pick one to three clear, consistent traits. Don't list their entire life story.
  • Check for agreement: If the group you are trying to mimic is split down the middle on what things mean, the robot probably won't be able to help.
  • It depends on the robot: Some robots (like Llama or Qwen) got better at these tasks with the right hints. Others (like DeepSeek) actually got worse, no matter what you told them.

In short, the researchers suggest that giving a robot a "persona" isn't a magic button that fixes everything. It's more like tuning a radio: you have to find the right frequency (the right few attributes) to get a clear signal. If you turn the dial too far or try to listen to too many stations at once, you just get static.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →