False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs
This paper challenges the view that high-confidence errors in large language models are merely fragile failures by demonstrating that they can be stable, internally coherent "false fixed points" where robustness diverges from truth-tracking, a phenomenon driven by mechanisms like high signal-to-noise inertia and representational compression.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: When Being "Stuck" Doesn't Mean Being "Right"
Imagine you are walking through a foggy forest. You come to a fork in the road. You are 100% sure you should go left, so you start walking confidently.
Usually, we think that if you are wrong, it's because you are "wobbly." If someone gently nudged you or asked, "Are you sure?", you would realize your mistake and change your mind. We assume wrong answers are fragile.
This paper challenges that idea.
The authors suggest that sometimes, an AI can be confidently wrong and also stubbornly stable. They call these "False Fixed Points." The AI isn't wobbling; it's locked into a wrong answer so tightly that even if you poke it or shake it, it doesn't budge. It's not a "brittle" mistake; it's a "solid" mistake.
The Core Metaphor: The Locked Door vs. The Wobbly Gate
Think of an AI's decision-making like a door.
- Fragile Error (The Wobbly Gate): The gate is loose. If you push it, it swings open, and you realize you went the wrong way. This is what we usually expect from a mistake.
- False Fixed Point (The Locked Door): The door is heavy, bolted shut, and made of steel. You push it, and it doesn't move at all. You are still on the wrong side of the forest, but the door is so stable that you never realize you need to find a different exit.
The paper argues that in Large Language Models (LLMs), these "Locked Doors" (stable wrong answers) are a real problem. The model isn't failing because it's confused; it's failing because it's too confident in a wrong path, and that confidence is surprisingly hard to break.
The "Kantian" Check-In (The Bouncer at the Club)
To fix this, the authors borrow an idea from the philosopher Immanuel Kant. Kant asked: "Do you actually have the right to make this judgment?"
Imagine a bouncer at a club (the "Commitment Gate").
- Normal AI: Walks up and says, "I'm going in!" (Even if it doesn't have a ticket).
- Kantian AI: The bouncer stops it and asks: "Do you have a ticket? Do you have evidence? Is the question even answerable?"
The paper tests a version of this where the AI is forced to pause and ask itself: "Am I allowed to commit to this answer, or should I just say 'I don't know'?"
What They Actually Found (The Experiments)
The researchers tested three different AI models with a bunch of true/false questions. Here is what happened:
1. The "Poke Test" (Are wrong answers wobbly?)
They took the AI's "Locked Doors" (confidently wrong answers) and the "Open Gates" (confidently correct answers) and gave them a gentle digital "poke" (a tiny change to the input).
- The Expectation: They thought the wrong answers would wobble and fall apart easily.
- The Reality: The wrong answers were just as solid and stable as the right ones. You couldn't tell the difference just by shaking them. This proves that stability does not equal truth. A model can be rock-solid and still be completely wrong.
2. The "I Don't Know" Strategy (The Trade-off)
They tried the "Bouncer" approach again, telling the AI: "If you aren't 100% sure, just say 'I don't know' instead of guessing."
- The Result: This worked! The AI made far fewer "confidently wrong" mistakes.
- The Catch: To do this, the AI had to stop answering questions it could have guessed on. It traded quantity (answering more things) for quality (being less wrong). It became safer, but it also became quieter.
3. The "Strict Gate" (C3-R)
They tried an even stricter version where the AI had to list specific reasons why it shouldn't answer before it could commit.
- The Result: This was the safest option. It almost eliminated confident mistakes. But, like the previous step, it meant the AI answered fewer questions overall. It was a very cautious, conservative guard.
Why Does This Happen? (The "Heavy Armor" Theory)
The paper offers a guess as to why these "Locked Doors" exist, using a physics analogy.
Imagine the AI's brain is a giant, heavy ship moving through water.
- High Signal-to-Noise (High-SNR) Inertia: Some models (like the Qwen2.5 they tested) have "heavy armor." Their internal signals are so massive and loud that tiny little pokes (perturbations) just bounce off them. The ship is so heavy that a small wave doesn't change its course.
- The Problem: This makes the ship very stable, but if the ship is steering toward a rock, that heavy armor makes it impossible to turn away from the rock just by giving it a little nudge. The "stability" is actually a trap.
The Bottom Line
This paper teaches us a simple but important lesson: Just because an AI sounds confident and doesn't change its mind when you poke it, doesn't mean it's right.
- Robustness (Stability) and Truth are two different things.
- We can't just look for "wobbly" answers to find mistakes; sometimes the worst mistakes are the steadiest ones.
- The best way to handle this right now is to teach the AI to say "I don't know" more often, even if that means it answers fewer questions.
The authors are careful to say this is a new way of looking at the problem, not a magic fix. They suggest we need new tools to measure "stability" and "truth" separately, rather than assuming they go hand-in-hand.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.