← Latest papers
💬 NLP

Coherence Maximization Improves Pluralistic Alignment

This paper demonstrates that maximizing internal coherence in AI-generated examples, rather than relying solely on label accuracy or extensive human supervision, effectively steers models toward diverse human values and significantly improves generalization across various benchmarks.

Original authors: Taslim Mahbub, Yiding Pei, Shi Feng

Published 2026-06-03
📖 4 min read☕ Coffee break read

Original authors: Taslim Mahbub, Yiding Pei, Shi Feng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a very smart, but slightly confused, robot how to understand different groups of people. The robot has read almost everything on the internet, so it knows a lot, but it tends to give a "one-size-fits-all" answer to everyone. It doesn't quite get that a person from one country might have different values than a person from another, or that a Democrat might think differently than a Republican.

The problem is: How do you teach the robot these specific differences without hiring a human teacher to sit down and write out thousands of examples for every single group? That would take forever and cost a fortune.

This paper introduces a clever trick called Internal Coherence Maximization (ICM). Here is how it works, using some simple analogies:

The "Jigsaw Puzzle" Analogy

Think of a person's values (like their political views or cultural beliefs) as a giant jigsaw puzzle.

  • The Old Way: To teach the robot, humans would usually hand it a box of puzzle pieces that are already labeled with the correct picture (Gold Labels). This works great, but it requires humans to do all the labeling.
  • The New Way (ICM): Instead of giving the robot the labeled pieces, the researchers let the robot look at a pile of unlabeled pieces and try to figure out how they fit together on its own.

The robot asks itself: "If I believe this piece goes here, does it make sense with the piece next to it? Do all these pieces tell a consistent story?"

The researchers found that if the robot arranges the pieces so they fit together logically (high coherence), it learns the group's values almost as well as if a human had labeled them perfectly.

The Big Surprise: "Consistency Beats Perfection"

Here is the most interesting part of the discovery. The researchers tested two types of puzzle pieces:

  1. Perfect Pieces: Pieces that are 100% correct according to human facts, but might be a bit jumbled or inconsistent with each other.
  2. Coherent Pieces: Pieces that might have a few small errors, but they all tell a perfectly consistent story that fits together like a well-built wall.

The Result: The robot learned much better from the Coherent Pieces. Even if the individual pieces weren't 100% perfect, the fact that they made sense together helped the robot understand the group's values much better.

It's like trying to learn a new language. If you memorize 100 sentences that are grammatically perfect but contradict each other, you'll be confused. But if you learn 10 sentences that are slightly imperfect but tell a consistent story about how the language works, you'll actually understand the language better.

When the Robot Gets Stuck (The "Underrepresented" Problem)

The robot is really good at this puzzle game for groups it has seen a lot in its training data (like people from the US or UK). But for groups it hasn't seen much (like people from Nigeria or Indonesia), the robot's internal puzzle pieces don't fit together very well. It gets confused.

The researchers found a shortcut for these tricky cases. Instead of asking a human to label random questions, they asked the robot: "Where are you most unsure?"

  • The robot pointed to the specific questions where its internal puzzle pieces were clashing.
  • Humans then only had to answer those few, specific questions.
  • With just a tiny bit of human help on the hardest parts, the robot's understanding of those underrepresented groups improved dramatically.

The Bottom Line

This paper shows that to align AI with diverse human values, we don't need a human to check every single answer. Instead, we can let the AI organize its own knowledge to find what makes the most sense internally.

  • Consistency is key: A set of examples that tells a consistent story is more powerful than a set of examples that are individually perfect but messy.
  • Smart help: For groups the AI doesn't know well, we only need to give it a tiny nudge on the specific things it's confused about, rather than re-teaching it everything.

This approach allows us to build AI that respects different cultures and viewpoints without needing a massive army of human labelers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →