ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
This paper introduces ProofCouncil, an open-source LLM agent utilizing an author-critic architecture that achieved state-of-the-art performance in the FirstProof challenge by autonomously solving six out of ten real-world mathematical problems and demonstrating significant progress on a broader set of open problems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've handed a brand-new, super-smart robot a stack of the world's most baffling math puzzles. These aren't your standard "find the missing number" riddles; they are open problems that real human mathematicians have been scratching their heads over for years, with no one knowing the answer yet.
Enter ProofCouncil. Think of it not as a single robot, but as a tiny, high-stakes math workshop.
The Workshop Crew: Author, Critic, and the Squad
The heart of this system is a clever game of "hot potato" played by two main characters: the Author and the Critic.
- The Author is the dreamer. It's an AI that sits down and tries to write a perfect proof (a step-by-step mathematical argument) for a problem. It has a magical notebook where it scribbles ideas, tries out code, and even searches the web for clues.
- The Critic is the strict editor. It reads what the Author wrote and says, "Wait, this step doesn't make sense," or "You forgot to prove this part!" It doesn't just say "wrong"; it gives specific feedback on how to fix the gaps.
Here's the twist: The Author doesn't just write once. It writes, gets criticized, rewrites, gets criticized again, and keeps going. It's like a writer and an editor locked in a room, polishing a story until it's perfect.
But sometimes, the Author gets stuck. That's when they call in the Council and the Compute Node.
- The Council is a panel of other super-smart AI models (like a team of guest experts). The Author asks them, "Hey, I'm stuck on this specific part; what's a different way to look at it?"
- The Compute Node is the heavy lifter. If the problem needs a massive calculation or a specialized math software tool that the Author can't run alone, this node does the heavy lifting, running code and crunching numbers to check if a hunch is true.
Every few rounds, the Critic takes a "mental break" and starts fresh with a clean slate. This is crucial because if the Critic gets too used to the Author's mistakes, it might stop seeing them. A fresh pair of eyes ensures the proof is actually solid.
The Big Test: The FirstProof Challenge
The team put ProofCouncil to the ultimate test: a challenge called FirstProof. The rules were simple but brutal: solve 10 real, unsolved math problems in 24 hours, using only public tools.
The results? ProofCouncil was the star of the show. Out of the 10 problems, it managed to produce solutions for 6 of them that human referees judged to be correct (with only tiny, minor tweaks needed). This was the best performance of any team in the competition.
They also tried it on 30 other open problems sent in by real mathematicians. Of the 21 problems they got feedback on:
- 5 were judged to be completely correct solutions.
- 2 were seen as very promising, just waiting for a final check.
- 8 made useful progress, even if they didn't finish the job.
- 4 had no errors but didn't really move the needle.
- 2 misunderstood the question and solved an easier version of it instead.
Importantly, no expert said any of the outputs were mathematically wrong in a way that broke the logic. The main failure wasn't bad math; it was sometimes misreading the problem statement.
The Cost of Genius
Here's the catch: being this smart is expensive. Running ProofCouncil on a single problem cost about $350 in computer time and model calls. Compare that to a simpler, single-shot AI attempt that cost only $12 but solved fewer problems. The paper suggests that while the "team approach" works better, it's currently a luxury item. The authors admit they didn't try to make it cheaper yet, leaving that as a job for future researchers.
What It Didn't Do (And What It Didn't Claim)
It's important to know what this robot didn't do.
- It didn't solve all 10 problems in the challenge (it missed 4).
- It didn't produce a polished, ready-to-publish research paper. The mathematicians who reviewed the work noted the writing was dense and needed human editing to be readable by non-experts.
- It didn't "prove" that AI can solve any math problem. It showed that for some hard problems, a team of AIs working together is better than one AI working alone.
The Bottom Line
ProofCouncil suggests that if you want to solve a really hard, open math problem, you shouldn't just ask one AI to "go figure it out." Instead, you need a system where one AI writes, another tears it apart, a third offers fresh perspectives, and a fourth does the heavy math calculations.
In the world of the FirstProof challenge, this "Author-Critic-Council" team was the most successful strategy tested so far. It didn't conquer the entire mountain, but it climbed higher than anyone else, proving that when AI agents work together like a human research team, they can tackle problems that were previously thought to be out of reach.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.