← Latest papers
📊 statistics

Rethinking Intrinsic Dimension Estimation in Neural Representations

This paper challenges the validity of common intrinsic dimension estimators in neural representation analysis by demonstrating a critical discrepancy between theory and practice, revealing that these methods fail to track true underlying dimensions and instead reflect other confounding factors.

Original authors: Rickmer Schulte, David Rügamer

Published 2026-04-23
📖 5 min read🧠 Deep dive

Original authors: Rickmer Schulte, David Rügamer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Fake" Compass

Imagine you are exploring a vast, foggy forest (the world of Artificial Intelligence). Scientists have been using a special compass called the Intrinsic Dimension (ID) to figure out how "complex" the terrain is.

The theory goes like this: Even though the forest looks huge and 3D, maybe the actual path you need to walk is just a thin, winding trail (a low-dimensional "manifold"). If you can find that trail, you understand the data better.

For years, researchers have used this compass to map neural networks (AI brains). They noticed a strange pattern: As the AI gets deeper into its thinking process, the compass says the terrain gets "bigger" and more complex. They thought this meant the AI was creating more abstract, complex ideas as it went deeper.

This paper says: "Stop. Your compass is broken."

The authors, Rickmer Schulte and David Rügamer, prove that the compass isn't actually measuring the complexity of the terrain. It's just measuring how much the AI is stretching the map.


The Core Problem: The Stretchy Rubber Sheet

To understand why the compass is broken, imagine the neural network is a giant, stretchy rubber sheet.

  1. The Input: You put a crumpled piece of paper (the data) on the sheet.
  2. The Layers: As the data moves through the AI, it passes through layers of "stretchers."
  3. The Result: By the time the data reaches the end, the rubber sheet has been pulled so tight that every single point on the paper is now very far away from its neighbors.

The "Broken Compass" (The Estimators):
The tools scientists use to measure complexity (called TwoNN and MLE) work by looking at how far apart neighbors are.

  • The Trap: Because the rubber sheet has been stretched so much, the neighbors look incredibly far apart.
  • The False Conclusion: The compass thinks, "Wow, these points are so far apart! The space must be huge and complex!"
  • The Reality: The space isn't complex; it's just stretched. The "Intrinsic Dimension" (the true complexity) is actually shrinking or staying the same, but the distance between points is growing.

The Mathematical Proof:
The authors proved mathematically that neural networks act like "Lipschitz mappings." In plain English, this means they can squish or stretch things, but they cannot create new complexity out of thin air. If you start with a simple shape, you can't turn it into a more complex shape just by stretching it. Therefore, the true complexity cannot go up as the AI gets deeper.

But the compass says it goes up. Why? Because the compass is confused by the stretching.


The "Token" Surprise: The Library of Discrete Books

The paper also looked at Large Language Models (LLMs) like the ones that write this text.

  • The Analogy: Imagine a library where every book is a unique word.
  • The Discovery: In an LLM, the "data" is just a finite list of words (tokens). Even though the AI creates a "space" for them, it's like a library with a finite number of books.
  • The Twist: Mathematically, a finite collection of points has zero complexity. It's like a single dot.
  • The Conclusion: If you measure the "Intrinsic Dimension" of an LLM's internal thoughts, the true answer should be zero. But the broken compass still says it's high because it's just measuring how far apart the "books" are on the shelf.

What Is Actually Happening? (The Real Story)

If the compass is lying, what is the AI actually doing?

The authors found that the "stretching" is driven by Entropy (a measure of how spread out the data is).

  • The Metaphor: Imagine a group of friends standing in a circle (the input). As they walk through the AI, they start walking further and further apart from each other, spreading out across a massive field.
  • The Cause: The AI is trying to make sure every different sentence or image gets its own unique spot in this field so it doesn't get confused.
  • The Result: The "Intrinsic Dimension" tools are just measuring how far the friends have spread out, not how complex their conversation is.

The paper shows that the "Intrinsic Dimension" curve (which goes up and then down) looks almost exactly like the Entropy curve. They are measuring the same thing: The expansion of space.


The Takeaway: A New Map for the Future

1. The Old Map is Wrong:
We can no longer trust the "Intrinsic Dimension" numbers we've seen in papers for the last few years. When a paper says, "The AI's complexity increases in the middle layers," they are actually saying, "The AI is stretching the data apart."

2. The New Perspective:
Instead of asking "How complex is the shape?", we should ask, "How much is the AI spreading the data out?"

  • Early Layers: The AI stretches the data out to separate different ideas (High expansion).
  • Late Layers: The AI starts to group similar ideas back together to make a final decision (Lower expansion).

3. The Future:
Scientists need to stop using these old compasses. Instead, they should use Entropy (spread) or other metrics that actually tell us what the AI is doing. The "Intrinsic Dimension" isn't a measure of the AI's intelligence; it's just a measure of how much the AI is stretching the rubber sheet.

Summary in One Sentence

The paper reveals that the popular tools used to measure the "complexity" of AI brains are actually just measuring how much the AI is stretching its data apart, meaning all the previous conclusions about AI "growing more complex" as it thinks were likely an illusion caused by a broken measuring tape.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →