Benchmarking Large Language Model Rationality Using Measurement Axioms
This paper introduces a measurement-theoretic framework based on utility axioms to evaluate the rationality of Large Language Models, revealing that many models exhibit systematic violations of preference transitivity that undermine decision coherence, even as some patterns mirror human behavior.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a new assistant to help you make important decisions, like choosing between different investment plans or medical treatments. You don't just want an assistant who gets the "right" answer on a test; you want one who is consistent. If they tell you that Plan A is better than Plan B, and Plan B is better than Plan C, they should logically conclude that Plan A is better than Plan C. If they change their mind based on how you ask the question or what they had for lunch, they aren't a reliable decision-maker.
This paper is like a rigorous "logic test" for 20 different Large Language Models (LLMs)—the AI brains behind tools like chatbots. Instead of asking, "Did you get the math right?" the researchers asked, "Are your choices logically consistent?"
Here is the breakdown of their experiment using simple analogies:
1. The Test: The "Rock-Paper-Scissors" of Logic
In the real world, we expect choices to be transitive. If you prefer apples to bananas, and bananas to cherries, you should prefer apples to cherries. If you suddenly prefer cherries to apples, your preferences are "intransitive" (broken).
The researchers set up a game where the AI had to choose between different "gamble" tickets. Each ticket had a chance of winning money.
- Ticket A: 7/24 chance to win $5.00.
- Ticket B: 8/24 chance to win $4.75.
- Ticket C: 9/24 chance to win $4.50.
They asked the AI to compare these tickets in every possible combination. If the AI is rational, its choices should form a straight line. If it creates a loop (A > B > C > A), it has failed the logic test.
2. The Three Ways They Tried to Break the AI
The researchers didn't just ask the questions once. They tried to "break" the AI's logic by changing three things, like a scientist tweaking a recipe to see if the cake collapses:
The "Temperature" Knob (Randomness):
Imagine the AI is a chef. At "low temperature," the chef is strict and follows the recipe exactly. At "high temperature," the chef gets creative and adds random spices. The researchers turned this knob from very strict to very chaotic. They wanted to see if making the AI more "creative" made its logic fall apart.- The Result: Surprisingly, making the AI more chaotic (higher temperature) actually made its logic look better on the easiest test, but that was likely just because it started guessing randomly (like flipping a coin), which accidentally fits the rules of the test. It didn't mean the AI was actually thinking more clearly.
The "Memory" Game:
Imagine you are taking a test. In one version, you take each question in isolation. In the other, you are told, "Remember, you just chose Option A for the last 10 questions; now choose again."- The Result: Giving the AI a "memory" of its past choices made it worse. Instead of becoming more consistent, the AI got confused. Its logic became shaky, and it changed its mind more often depending on the order of the questions.
The "Wording" Trick:
They asked the exact same question but changed how the numbers looked.- Version 1: "7/24 chance" (a fraction).
- Version 2: "29.17% chance" (a percentage).
- The Result: The AI's logic changed just because the numbers looked different. If it was truly rational, 7/24 and 29.17% should be treated exactly the same. The fact that the AI treated them differently shows its "brain" is sensitive to the surface appearance of the data, not just the underlying truth.
3. The "Human-like" Mistake
Here is the most fascinating part. When the AI failed the logic test, it didn't fail randomly. It failed in a very specific way that humans also fail.
Decades ago, psychologists found that humans often use a "shortcut" rule: "If the difference in winning chance is tiny, I'll just pick the one with the most money. But if the difference in winning chance is huge, I'll pick the safer one." This shortcut causes humans to make illogical loops.
The AI models didn't just make random errors; they mimicked this specific human shortcut. However, they did it much more often than actual humans do. It's as if the AI is trying to be human, but it's overdoing the "human error" part.
4. The Big Conclusion
The paper concludes that while these AI models are great at generating text and answering trivia, they do not have a stable, internal "value system" like a rational human or a perfect computer program.
- They are not "rational agents" in the strict sense.
- Their choices depend heavily on how you ask the question, how "random" you tell them to be, and whether they remember what they just said.
- If you rely on an AI to make a series of connected decisions (like a doctor diagnosing a patient step-by-step), you cannot trust that its logic will hold up from step one to step ten.
In short: The AI is a brilliant mimic that can sound very smart, but if you poke it with a logic stick, it wobbles. It doesn't have a solid, unchanging set of beliefs; it just reacts to the immediate shape of the question.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.