← Latest papers
💬 NLP

Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning

This paper challenges the notion that representational convergence in large language models implies shared reasoning strategies, revealing that while models develop similar internal representations for shared inputs, they exhibit significant dissociations in reasoning processes, particularly regarding problem difficulty, decision stages, and the causal influence of shared information.

Original authors: Muhammad Usama, Dong Eui Chang

Published 2026-05-25
📖 5 min read🧠 Deep dive

Original authors: Muhammad Usama, Dong Eui Chang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a group of 16 different students (the AI models) sitting in a classroom, all trying to solve the same 800 difficult riddles. These students have different teachers, different textbooks, and even different ways of thinking.

A popular theory in the AI world, called the Platonic Representation Hypothesis, suggests that as these students get smarter, they all start thinking in the exact same way. It's like they are all discovering the same "universal truth" about how the world works.

This paper asks a simple question: If these students look at the problem in the same way, do they also solve it in the same way?

The answer is a surprising "No." The paper finds that while these models agree on what they see, they completely disagree on how they think. Here is the breakdown using simple analogies:

1. The "Confusion is Contagious" Effect (Difficulty Inversion)

You might expect that when students solve a hard problem correctly, they would all use similar, brilliant logic. You would expect them to look very similar when they are right.

What actually happens:

  • When they get it right: The students use very different, unique strategies. Their "thought processes" look nothing alike.
  • When they get it wrong: They all look exactly the same.

The Analogy: Imagine a group of people trying to navigate a dense fog.

  • When the path is clear (an easy problem), everyone takes a different, efficient route to the destination. They look different because they are successful.
  • When the fog is thick (a hard problem), everyone gets lost. They all end up standing in the same spot, staring blankly, because the fog makes everyone's vision blurry in the same way.
  • The Finding: The models are most similar when they are confused, not when they are smart. They converge on "shared confusion" rather than "shared understanding."

2. The "First Impression vs. Final Decision" Gap (Generation Gap)

The researchers looked at the models' brains at two different times:

  1. Before the decision: When they are just reading the question.
  2. After the decision: When they are actually generating the answer.

What actually happens:

  • Reading the question: All models look almost identical. They agree on what the words mean.
  • Making the answer: The moment they have to decide what to say, their brains diverge wildly. They go in completely different directions.

The Analogy: Think of a group of chefs looking at the same basket of ingredients.

  • Step 1 (Reading): They all agree on what the ingredients are (this is the "pre-decision" phase). They all see a tomato and an onion.
  • Step 2 (Cooking): Once they start cooking, one chef makes a salad, another makes a soup, and a third burns it. Their final dishes (the output) are totally different, even though they started with the same ingredients.
  • The Finding: The models agree on the input (the ingredients) but disagree on the computation (the cooking).

3. The "Ghost Knowledge" Effect (Epiphenomenal Correctness)

The researchers found that the models do store the correct answer in their shared "thoughts." If you asked a smart probe (a little detective) to look at Model A's brain, it could tell you the answer. If that detective then looked at Model B's brain, it could also tell you the answer.

However:

  • Just because the information is there doesn't mean the model uses it to make its decision.
  • If you tried to "delete" that specific piece of knowledge from the model's brain, the model would barely notice. It would still give the same answer.

The Analogy: Imagine a student taking a test who has the answer written on their hand, but they are too nervous to look at it.

  • The answer is physically present on their hand (shared representation).
  • But the student doesn't actually use that information to solve the problem; they guess instead.
  • The Finding: The models carry the "truth" in their minds, but it doesn't actually drive their behavior. It's like a passenger in a car who knows the destination but isn't the one driving.

Why Does This Happen?

The paper suggests a mechanism involving Attention.

  • When a problem is easy, the model's "attention" (focus) is sharp and specific. Different models focus on different details, leading to different solutions.
  • When a problem is hard, the model's attention becomes "diffuse" or blurry. It spreads out over everything, like a flashlight beam that is too wide. This blurriness makes all models look the same because they are all just "guessing" in a similar, unstructured way.

The Bottom Line

The paper concludes that AI models are not converging on a shared "logic" or "truth." Instead, they are converging on a shared way of processing input (like how we all see a red apple as red).

  • What they share: How to read the prompt.
  • What they don't share: How to reason through the problem.

This means that if you want to build a team of AI models that work well together (an "ensemble"), you shouldn't pick them because they look different on paper. You should pick them because they make different mistakes or use different strategies when solving hard problems, because that is where their true differences lie.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →