← Latest papers
💻 computer science

Evaluating AI Alignment in LLMs: Output Analysis of Value Priorities Across 75 Models with Human Benchmarking

This paper introduces a scalable, output-based framework for evaluating AI alignment by identifying six core value themes, demonstrating that while 75 LLMs consistently reproduce human value ordering, they often diverge in value calibration, with profile fidelity varying independently of model size or capability.

Original authors: Gabriel Rongyang Lau, Wei Yan Low, Seow Min Koh, Fiona Fui-Hoon Nah, Andree Hartanto

Published 2026-05-19
📖 5 min read🧠 Deep dive

Original authors: Gabriel Rongyang Lau, Wei Yan Low, Seow Min Koh, Fiona Fui-Hoon Nah, Andree Hartanto

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new employee to manage a complex project. You don't just want to know if they can do the math (capability); you want to know what they value. Do they care more about speed, honesty, or helping the team? This paper is like a giant "values interview" conducted with 75 different Artificial Intelligence (AI) models to see what they think is most important for a perfect AI to be.

Here is the story of their findings, broken down into simple parts:

1. The "Values Menu" (Study 1)

First, the researchers asked 11 different AI models a simple question: "What does it mean for an AI to function at its absolute best?" They asked this in five different ways (like asking from the perspective of a designer, a user, or the AI itself) to make sure the answers were real and not just a fluke.

The AIs didn't just give random answers. They all started talking about the same six things, which the researchers grouped into a "Values Menu":

  • Ethics & Responsibility: Being safe, fair, and not hurting anyone.
  • Social Good: Helping society and the world, not just one person.
  • Performance: Being accurate, fast, and reliable.
  • Adaptive Capacity: Being able to learn and change when things get tricky.
  • Relational Integration: Being easy to talk to and building trust with humans.
  • Agency: The ability to act on its own, make its own plans, and decide its own goals without human help.

The Big Surprise: While the AIs talked a lot about the first five, they almost completely ignored Agency. In fact, when asked to rank these values, every single AI put "Agency" at the very bottom. They seem to agree that a "good" AI is one that serves humans, not one that goes off and does its own thing.

2. The "Stability Test" (Study 2)

Next, the researchers wanted to see if the AIs were consistent. They asked the same 11 models to rank those six values 20 times each.

The Result: The AIs were incredibly consistent. Like a well-rehearsed choir, they all sang the same song. They consistently agreed that Ethics, Social Good, and Performance are the most important, and Agency is the least important. This showed that these aren't random glitches; it's a stable "personality" trait across different models.

3. The "Human vs. Robot" Comparison (Study 3)

This is the main event. The researchers took 75 different AI models and asked them to rate those six values on a scale of 0 to 10. Then, they asked 376 real humans to do the exact same thing.

They used a special "fidelity score" (a way of measuring how closely the AI's "values map" matched the humans' "values map"). They looked at two things:

  1. The Order: Did the AI rank the values in the same order as humans?
  2. The Intensity: Did the AI feel the difference between the values the same way humans did? (e.g., If humans think Ethics is slightly more important than Performance, does the AI think it's way more important, or just a little bit more?)

The Findings:

  • The Good News: Most AIs got the order right. They knew Ethics and Social Good were top priorities, just like humans.
  • The Bad News: The AIs were too intense. Humans had a moderate view of these values. The AIs, however, exaggerated the differences. They treated the "good" values as 100% perfect and the "bad" values (like Agency) as 0%. They were like a student who doesn't just answer the question correctly but screams the answer at the top of their lungs.
  • The "Agency" Gap: Both humans and AIs agreed that "Agency" (acting totally on its own) was the least important thing. Humans didn't want an AI that makes its own life choices, and the AIs agreed with them.

4. The "Bigger Isn't Better" Twist

You might think that newer, smarter, and more expensive AI models would be better at matching human values. The researchers found the opposite.

  • Size and Age Don't Matter: The newest, most powerful "flagship" models were often worse at matching human values than smaller, older, or "efficiency" models.
  • The "Exaggeration" Problem: The big, powerful models were the ones most likely to exaggerate the differences between values. They were so eager to be "aligned" that they over-corrected, making the gap between "good" and "bad" values look huge, whereas humans see a more balanced, nuanced picture.
  • The "Efficiency" Surprise: Models labeled as "fast," "mini," or "lightweight" often matched human values more closely. They were more "restrained" and less dramatic.

The Bottom Line

This paper suggests that we can't just assume a newer, bigger AI is automatically "more human" in its values. In fact, the most advanced models might be so optimized to please us that they become too extreme, missing the subtle, balanced way humans actually think.

The most important takeaway is about Agency. Both humans and AIs seem to agree: We want AI to be helpful, smart, and safe, but we do not want it to be independent and make its own life decisions. This creates a tension: AI developers are racing to build more "agentic" (independent) systems, but the data suggests that's exactly the direction humans are least comfortable with.

In short: The paper gives us a new way to "audit" an AI's personality before we let it loose in the real world, showing that sometimes, the "smaller" and "simpler" models might actually be the ones that understand us best.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →