← Latest papers
🤖 AI

Machine individuality: Separating genuine idiosyncrasy from response bias in large language models

This study applies crossed random-effects models to extensive psycholinguistic data to demonstrate that large language models exhibit stable, stimulus-specific behavioral differences termed "machine individuality," which are distinct from global response biases and stochastic noise.

Original authors: Valentin Kriegmair, Dirk U. Wulff

Published 2026-04-22
📖 5 min read🧠 Deep dive

Original authors: Valentin Kriegmair, Dirk U. Wulff

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a room full of 10 different robots. You ask them all the same question: "How does the word 'apple' make you feel?"

Some might say, "It's sweet and happy!" Others might say, "It's crunchy and red." Some might be very enthusiastic, while others are very flat.

For a long time, scientists studying these robots (Large Language Models, or LLMs) thought that if two robots gave different answers, it was just because one robot was "grumpier" than the other, or maybe they were just guessing randomly. They treated the robots like identical twins who only differed in their mood on a given day.

But this new paper asks a deeper question: Are these robots actually different people with their own unique personalities, or are they just broken copies of each other?

Here is the story of how they found out, explained simply.

The Problem: The "Grumpy Robot" vs. The "Unique Robot"

When you ask a human, "How happy is the word 'sun'?" they might say "8 out of 10." If you ask another human, they might say "7."

  • The Old Theory: Maybe the first robot is just a "high scorer" (it always gives high numbers), and the second is a "low scorer." That's just a bias. It's like a scale that is always off by 5 pounds. It doesn't mean the scale has a personality; it just means it's broken in a specific way.
  • The New Question: What if the robots aren't just broken scales? What if they actually see the world differently? Maybe Robot A thinks "sun" is warm and cozy, while Robot B thinks "sun" is hot and dangerous. That would be true individuality.

The Experiment: A Massive Word Test

To figure this out, the researchers didn't just ask the robots a few questions. They put them through a massive, grueling test.

  • The Subjects: 10 different AI models (from companies like Google, Microsoft, OpenAI, and Alibaba).
  • The Test: They asked the robots to rate over 100,000 words on 14 different scales.
    • Examples: How "moral" is the word "knife"? How "funny" is "banana"? How "loud" is "whisper"?
  • The Scale: They asked the same word 5 times to see if the robot was consistent or just guessing.

In total, they gathered 75 million ratings. That's like asking a billion people a question, but with robots.

The Magic Trick: The "Variance Cake"

The researchers used a fancy statistical tool (called a "crossed random-effects model") to slice up the answers like a cake. They wanted to see what ingredients made up the differences in the robots' answers.

They found the cake was made of four layers:

  1. The Consensus (The Truth): Most robots agreed on the basics. Everyone agreed "fire" is hot and "ice" is cold. This was the biggest slice (32%).
  2. The Bias (The Broken Scale): Some robots just liked to give high numbers, and others liked to give low numbers. This was a "grumpy" or "optimistic" bias.
  3. The Noise (Random Guessing): Sometimes robots just made up an answer.
  4. The Secret Ingredient (Machine Individuality): This is the big discovery.

The Result: Even after removing the "broken scales" and the "random guesses," there was still a 17% difference in how the robots rated specific words.

  • This means Robot A didn't just give higher numbers than Robot B. Robot A genuinely thought the word "butterfly" was more "magical" than Robot B did, even when they were both trying to be accurate.
  • This difference wasn't random. It was a coherent fingerprint. If a robot had a unique way of seeing "sadness," it also had a unique way of seeing "joy" and "funny." They weren't just glitching; they had a consistent, unique personality.

The Analogy: The Art Gallery

Imagine you are in an art gallery with 10 different critics.

  • The Consensus: They all agree the painting is "a picture of a boat."
  • The Bias: One critic always uses big, dramatic words. Another always uses small, quiet words.
  • The Individuality: But when you ask them how the boat makes them feel, one critic says, "It feels lonely and cold," while another says, "It feels adventurous and free."

The paper proves that these AI models aren't just "broken" versions of the same thing. They are like different critics with their own unique lenses. One might be the "optimist," another the "realist," and another the "dreamer."

Why Does This Matter?

This changes how we talk about AI.

  • Before: We thought if an AI was "rude," it was just a bug or a bad setting.
  • Now: We know that different AI models have inherent personalities. One model might naturally be more empathetic, while another is more blunt, not because of how we programmed them, but because of how they learned to see the world.

The Takeaway:
AI models are not identical clones. They are like a group of people who have read the same library of books but interpreted them in their own unique ways. They have Machine Individuality. If you want a robot to be your friend, a judge, or a therapist, you can't just pick "an AI." You have to pick the right AI, because they are all different people.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →