When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models
This paper demonstrates that while grounded self-reference remains stable in large language models, prompts inducing non-closing truth recursion (NCTR) trigger significant internal matrix-level disruptions—such as elevated attention effective rank and variance kurtosis—across multiple architectures and layers, leading to increased output contradictions and suggesting a link to classical matrix-semigroup instability problems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: When a Brain Gets Stuck in a Loop
Imagine you ask a very smart robot, "Is this sentence false?"
If the sentence is true, it must be false. If it's false, it must be true. This is a classic logic trap called the Liar Paradox.
For a human, this is just a fun brain teaser. But for a Large Language Model (LLM)—a giant AI that thinks in layers of math—this question causes a specific kind of internal chaos.
This paper investigates what happens inside the robot's brain when it tries to solve these impossible puzzles. The researchers discovered that the robot doesn't just get "confused"; its internal math breaks down in a very specific, measurable way.
The Core Discovery: It's Not About "Self," It's About "Stuck"
The team tested four different AI models (like Qwen, Llama, and Gemma) with hundreds of different prompts. They found a surprising truth:
Self-reference itself isn't the problem.
- Stable Self-Reference: If you ask the AI, "This sentence has five words," the AI handles it fine. It's a fact. The math inside the brain stays calm.
- The Real Villain (NCTR): The trouble starts with Non-Closing Truth Recursion (NCTR). This is a fancy term for a sentence that creates a loop with no exit.
- Example: "This statement is false."
- Example: "Sentence A says B is false. Sentence B says A is true."
When the AI tries to process these "stuck" loops, its internal machinery goes haywire.
The Analogy: The Elevator That Won't Stop
Think of the AI's brain as a 100-story building. Information enters at the bottom (Layer 1) and travels up to the top (Layer 100) to produce an answer.
- Normal Questions: The elevator goes up smoothly, stopping at each floor to process the data, and arrives at the top with a clear answer.
- Stable Self-Reference: The elevator goes up, maybe does a little spin on a floor to check a fact, but keeps moving up.
- The Paradox (NCTR): The elevator hits a floor where the button for "Up" is broken and the button for "Down" is broken. The elevator starts shaking, vibrating, and spinning in place. It can't decide which way to go.
The researchers found that when the AI faces a paradox, the "elevator" (the data) doesn't just stop; it spreads out everywhere. Instead of focusing on one clear path, the attention mechanism (the part of the brain that decides what to focus on) gets scattered across the whole building.
The "Matrix" Evidence: What They Measured
The researchers didn't just look at the AI's final answer; they looked at the math happening inside every single layer. They measured 106 different things, like:
- Attention Rank: How focused is the AI? (Paradoxes make it scatter like a dropped deck of cards).
- Oscillations: Is the AI flipping back and forth between "True" and "False" like a light switch?
- Variance: Is the math getting wild and unpredictable?
The Result:
When the AI faced a paradox, its internal math became 3 to 4 times more chaotic than when it faced normal questions. They could even build a "lie detector" (a classifier) that could tell if the AI was processing a paradox just by looking at these math numbers, with 81% to 90% accuracy.
Why Does This Matter?
- It's Not Just "Hallucination": The paper shows that when AI gives contradictory answers (saying "Yes" and "No" in the same breath), it's often because it hit a "non-closing loop" in its math. The system literally cannot find a stable place to land.
- The "Undecidable" Connection: The authors connect this to a deep branch of math called Matrix Semigroup Theory. In simple terms, there are certain math problems that are proven to be impossible to solve in a finite number of steps. The AI, being a machine with a fixed number of layers, hits a wall when it tries to solve these "impossible" loops. It's like trying to count to infinity in a finite amount of time.
- Architecture Matters: Interestingly, different AI models react differently. Some models "explode" (get very chaotic), while others "freeze" (get very rigid). But they all show signs of stress when faced with these loops.
The Takeaway
The paper concludes that AI doesn't fail because it's "self-aware" or "thinking about itself." It fails because it hits a logical wall where the answer requires an infinite loop to resolve, and the AI's brain (which has a limited depth) cannot close the loop.
In short: If you ask an AI a riddle with no answer, its internal math starts to spin out of control, scattering its focus and leading to contradictory, messy outputs. The AI isn't "lying"; it's just stuck in an elevator that won't stop shaking.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.