The Undecidability of Artificial General Intelligence (AGI) Alignment
This paper establishes that AGI alignment is structurally unverifiable rather than impossible, proving through Trakhtenbrot's Wall and a derived Soundness-Completeness-Tractability Trilemma that current containment strategies are not temporary fixes but necessary sacrifices of logical expressivity to achieve decidable safety.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: You Can't Prove a Robot is Safe
Imagine you are building a super-smart robot (an AGI) that can learn anything and solve any problem. Your goal is to write a "safety manual" that proves, 100% of the time, that this robot will never hurt anyone or go rogue.
This paper argues that it is mathematically impossible to write such a perfect safety manual.
The author isn't saying the robot will go rogue. The argument is that you can never prove it won't. It's not a problem of bad engineering or slow computers; it's a fundamental law of logic, like gravity. No matter how smart your robot is, or how much money you spend on testing, you cannot create a universal "safety certificate" that works for every possible situation.
The Three Walls You Can't Climb
The paper says there are three ways people try to prove a robot is safe, and the author shows that all three hit a "wall" where logic breaks down.
1. The Infinite Wall (The "Everything" Problem)
The Idea: You try to test the robot in every possible situation, forever.
The Analogy: Imagine trying to check every single sentence a person could ever say to ensure they never tell a lie.
The Problem: Because the robot is smart enough to think about itself (like a human), it can create complex loops and riddles. The paper uses Rice's Theorem and Gödel's Incompleteness to show that if a system is smart enough to do general math, there will always be some "safe" behavior that looks dangerous, or some "dangerous" behavior that looks safe, and you can never write a rule to tell the difference. It's like trying to catch a shadow with a net; the more you try to define it, the more it slips away.
2. The Finite Wall (The "Hardware" Problem)
The Idea: "Okay, let's stop thinking about infinity. The real world is finite. The robot has a battery, a processor, and a limited amount of memory. If we just check every single thing it can do on this specific hardware, we can prove it's safe, right?"
The Analogy: Imagine a chessboard. It's finite (64 squares). You could, in theory, calculate every possible move.
The Problem: The paper introduces Trakhtenbrot's Wall. It says that while you can check one specific chessboard, you cannot write a single rule that guarantees safety for every possible computer configuration in the universe.
If you try to make a "Universal Safety Rule" that works for any finite computer (any size, any chip), the math says the rule itself becomes impossible to compute. It's like trying to write a single instruction manual that works for every possible car ever built; the manual would have to be infinitely long and complex, making it impossible to read or verify.
3. The Complexity Wall (The "Chess" Problem)
The Idea: "What if we just check one specific computer, one specific time, and one specific set of rules? Can't we just brute-force it?"
The Analogy: Think of a chess game. We know the game is finite. We know there is a "perfect" move for every situation. But the number of possible games is so huge (more than the number of atoms in the universe) that even if you had a computer the size of the universe, it would take longer than the lifespan of the universe to calculate the perfect move.
The Problem: The paper argues that checking a smart robot is like this. Even if the robot is trapped in a small, finite box, the number of ways it can behave is so complex that checking them all is intractable. It's not "impossible" in theory, but it's impossible in practice. It requires more computing power than exists in the universe.
The "Trilemma": You Can Only Have Two Out of Three
The paper concludes with a "Trilemma." Imagine you want three things from your safety system:
- Soundness: It never gives a false "Safe" signal (it's never wrong when it says "Go").
- Completeness: It never misses a danger (it catches every single bad thing).
- Tractability: It gives you the answer quickly (in a reasonable amount of time).
The Paper's Verdict: You can have two, but never all three.
- If you want it to be Fast and Never Wrong, you have to accept that it will miss some dangers (it's incomplete).
- If you want it to be Never Wrong and Catch Everything, it will take forever to give you an answer (it's intractable).
- If you want it to be Fast and Catch Everything, it will sometimes lie and say a dangerous robot is safe (it's unsound).
What This Means for Engineers
The paper says that current engineers are already doing the only thing they can: Sacrificing Completeness.
To keep robots safe, we deliberately make them "dumber" or "blinder." We do this by:
- Shielding: Putting the robot in a cage where it can only speak a simple language (so it can't ask complex, dangerous questions).
- Proof-Carrying Code: Forcing the robot to show a math proof before it acts (which means it can't do anything complex that it can't prove instantly).
- Short Horizons: Telling the robot, "Only think 5 seconds into the future," so it can't plan long-term tricks.
The Final Takeaway
The paper's main message is a bit sobering: The core barrier to AI safety isn't that we can't build a safe robot; it's that we can never mathematically prove it is safe.
The "safety" we have today is a temporary illusion created by restricting the robot's freedom. We force the robot to operate in a tiny, simple box so we can check it. But the moment we let the robot be truly "General" (smart enough to do anything), the math says we lose the ability to verify it. We are trading the robot's full potential for our peace of mind.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.