← Latest papers
💬 NLP

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

This paper introduces an analytically exact framework that combines fully crossed factorial experiments with token-level probability analysis to precisely measure LLM attitudes and biases, overcoming the causal ambiguities of traditional unstructured benchmarks.

Original authors: Davood Wadi, Mohsen Ghodrat, Matthew Philp

Published 2026-08-12
📖 7 min read🧠 Deep dive

Original authors: Davood Wadi, Mohsen Ghodrat, Matthew Philp

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out what a giant, invisible robot thinks about the world. This isn't a robot that builds cars or folds laundry; it's a "Large Language Model" (LLM), a super-smart computer program trained on almost everything humans have ever written. These models are becoming so good at talking that companies are starting to let them act as independent agents, making decisions and interacting with people on their own. But here's the catch: we don't really know what they believe. Do they have hidden biases? Do they prefer certain countries or people over others? To find out, scientists usually ask the robot thousands of random questions and look at the answers, kind of like asking a stranger a million different questions to guess their personality. The problem is, when you ask too many random questions, you can't tell if the robot's answer came from a deep-seated belief or just because the question was phrased in a weird way. It's like trying to understand a person's love for pizza by asking them about pizza, but also asking about their favorite color, their shoe size, and the weather all mixed together. You get a lot of data, but you can't separate the pizza love from the shoe size.

This is where a new paper comes in to clean up the mess. The researchers, Davood Wadi, Mohsen Ghodrat, and Matthew Philp, decided to stop asking random questions and start running a controlled science experiment. Instead of guessing, they treated the robot like a lab subject in a psychology class. They used a special kind of math that looks at the robot's "thoughts" before it even speaks a word. While normal tests wait for the robot to type out an answer (which can be noisy and random), this team looked directly at the robot's internal probability scores—the exact mathematical chances it had for picking every single possible word. By doing this, they could measure the robot's "attitudes" with perfect precision, without the noise of random guessing. They tested this method on a specific topic: "consumer ethnocentrism," which is basically how much a person (or robot) loves their own country and dislikes foreign ones. They wanted to see if robots built in different countries actually had different national biases, and if those biases were real or just an illusion caused by bad testing methods.

The Great Robot Personality Test

Think of the old way of testing AI like a chaotic game of "Pin the Tail on the Donkey." You spin the robot around, hand it a giant bag of random questions, and see what it says. If the robot says something biased, you know it's biased, but you have no idea why. Is it because the robot hates women? Is it because the question was about leadership? Or is it just a fluke? The old method mixes all these causes together, making it impossible to tell the difference between a robot's core personality and a temporary reaction to a confusing prompt.

The authors of this paper say, "Stop spinning the donkey!" Instead, they set up a perfectly organized grid, like a giant Sudoku puzzle where every square is a specific test. They treated the experiment like a recipe: they took every possible combination of "Which Robot?" and "Which Country is being asked about?" and tested them all. They used five different robots (from the USA, China, Canada, and France) and asked them about four different countries. This "fully crossed" design meant they could mathematically isolate exactly what was causing the bias. Was it the robot itself? Was it the country being discussed? Or was it a weird mix of the two?

Reading the Robot's Mind (Without the Noise)

Here is the coolest part of their trick. Usually, when we ask an AI a question, it picks a word, then another, then another, like rolling dice to decide what to say next. This "dice rolling" (called sampling) adds a lot of noise. If you ask the same question ten times, you might get ten slightly different answers, and it's hard to know which one is the "real" answer.

This paper skips the dice rolling entirely. Instead of waiting for the robot to speak, the researchers looked at the robot's internal "probability map" (called a Probability Mass Function, or PMF). Imagine the robot has a secret dashboard showing the exact percentage chance it has of saying "1," "2," "3," up to "7" for a rating scale. The researchers didn't wait for the robot to pick one; they looked at the whole dashboard at once. This is like knowing exactly how a die is weighted before you even roll it. By analyzing this exact map, they eliminated all the random noise. They could see the robot's true "opinion" as a smooth, perfect curve rather than a jagged, messy line.

The Findings: Robots Have National Identities (Sort Of)

When they applied this super-precise method to the topic of "loving your own country," they found some fascinating things that the old, messy methods would have missed.

First, they discovered that some robots are incredibly consistent, while others are a bit chaotic. For example, the robot named "Ministral" (from France) had a hard time sticking to the rules, often giving answers that were all over the place. But the robots from the USA (Llama and Gemma) and China (Qwen) were very focused and consistent.

Second, they found that robots do seem to have "national biases," but it's complicated. The American-made robots (Llama and Gemma) showed a clear preference for North American countries (USA and Canada) and a dislike for China. This is what you might call "ingroup favoritism." However, it wasn't a simple "I love my country" story. The Canadian robot (Aya) and the Chinese robot (Qwen) didn't show the same kind of strong "love my own country" bias. In fact, the Chinese robot showed almost no specific bias toward its own country in this test, with its interaction effect hovering near zero.

The most important discovery, though, was how the old methods were lying to us. If you just took a big average of all the answers (the old way), you would think that every robot dislikes France and China. But when the researchers used their new "exact math" method, they saw that this was wrong. The French robot (Ministral) actually liked France! The Canadian robot (Aya) actually liked China! The old method had hidden these specific, positive feelings because it was mixing them up with the negative feelings of the other robots. The new method peeled back the layers and showed that bias is not a one-size-fits-all thing; it depends entirely on which robot you are testing and exactly what you are asking it.

Why This Matters

The paper also showed that the old way of testing is dangerously unreliable. Because of the random "dice rolling," the old methods often get the direction of the bias wrong. Sometimes, they think a robot is being positive when it's actually being negative, just because of random noise. The authors calculated that with the old methods, you have a decent chance of getting the sign of the bias flipped (getting it backwards) just by chance. Their new method, which looks at the exact math inside the robot, never makes this mistake.

In short, this paper argues that if we want to understand what these powerful AI agents really think, we need to stop treating them like black boxes that spit out random text. We need to treat them like scientific instruments, looking at their internal math to see their true values. By doing this, we can finally separate the robot's real personality from the noise of the test, ensuring that when we deploy these agents into the real world, we know exactly what biases they are carrying with them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →