From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
This position paper argues for a paradigm shift in AI for Mathematics from solving predefined problems to acting as research agents capable of tackling frontier challenges, while providing a systematic review of current limitations and outlining a strategic roadmap for future development.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of mathematics as a giant, endless library. For a long time, the smartest AI "students" in this library were like brilliant Solvers. They were amazing at taking a specific, well-written homework assignment, reading the instructions, and finding the answer. If you gave them a math problem from a high school competition, they could often solve it faster than a human genius. They were like race cars on a perfectly paved track: fast, precise, and reliable.
But the authors of this paper argue that we are hitting a wall. These Solvers are great at finishing the race, but they can't invent the race. They can't look at the library, find a dusty, unsolved mystery that has been sitting there for decades, and figure out a brand-new way to crack it.
The paper suggests it's time to stop building better race cars and start building Research Agents. These aren't just students who follow instructions; they are explorers who can wander into the unknown, get lost, try weird ideas, and maybe—just maybe—discover something new.
The Current State: The "Solver" Era
Right now, AI is crushing it on known challenges.
- The MiniF2F Benchmark: This is like a super-hard high school math test. In 2021, AI got about 30% right. By July 2025, a system called Seed-Prover got 99.6% (243 out of 244 problems) correct. It's basically a perfect score.
- The IMO (International Mathematical Olympiad): These are the hardest problems for the world's best high schoolers. In 2024, an AI called AlphaProof got a "Silver Medal" level score. By July 2025, Seed-Prover solved 5 out of 6 problems from the 2025 competition.
This is huge! It proves that AI can do rigorous, machine-checked math. But here's the catch: these are all problems where the answer already exists in the universe of math. The AI is just finding the path to a known treasure.
The Problem: The "Research Gap"
The paper points out that real mathematical research is messy. It's not a multiple-choice test. It's like trying to find a needle in a haystack, but you don't even know what the needle looks like yet.
The authors argue that current AI systems are fundamentally limited when it comes to "frontier research."
- The "Rediscovery" Trap: When AI claims to solve a famous open problem (like one from the Erdős Problems list, a collection of over 1,100 tough math riddles), it often just finds a solution that humans already wrote down in a book, but the AI didn't know about.
- The "Hallucination" Risk: If you ask an AI to solve a problem it doesn't know the answer to, it might just make up a proof that looks right but is actually nonsense. Unlike a human who can say, "I'm stuck, I need a new idea," the AI often confidently walks off a cliff.
The paper explicitly rules out the idea that current AI is ready to be a solo mathematician discovering the next big theorem. They say systems like DeepSeek-Prover or AlphaProof are still just "Solvers," not "Researchers." They can't yet tackle the truly unsolved giants like the Riemann Hypothesis or the Navier-Stokes equations without human help.
The Roadmap: From Solver to Researcher
So, how do we turn a race car into an explorer? The authors lay out a map with five big hurdles to clear:
1. The Data Desert
Imagine trying to learn to write a novel by reading only 10 pages of text. That's the problem with formal math. While there are billions of words of normal text on the internet, there are only about 1.3 million lines of formal math code (in a system called Lean).
- The Fix: We need to teach AI to translate "human math" (which is vague and relies on context) into "computer math" (which is super precise). This is called Autoformalization. The paper suggests this is getting better, but it's still tricky. A tiny mistake in translation can make a hard problem look easy, or an easy problem look impossible.
2. The Missing Map
Current AI treats every math problem as a separate island. But in real math, everything is connected. A proof in geometry might rely on a lemma from algebra.
- The Fix: We need to build a giant Knowledge Graph. Imagine a web where every theorem is a node, and the lines connecting them show how they relate. If the AI can see these connections, it can use old ideas to solve new problems, rather than starting from scratch every time.
3. From Checking to Discovering
Right now, AI is great at checking if a proof is right (Verification). It's bad at coming up with the proof in the first place (Discovery).
- The Fix: We need systems that can conjecture (guess) new theorems and then try to prove them. Some experiments, like AlphaEvolve, are trying to do this by letting the AI "evolve" new ideas, similar to how nature evolves species. But the paper notes this is still in the early stages; the AI can't yet invent a new mathematical concept (like a new type of number) the way a human genius can.
4. The Toolbelt
Human mathematicians don't just use their brains; they use calculators, computers, and software to do the heavy lifting.
- The Fix: AI needs to learn to use these tools too. But there's a catch: if the AI asks a calculator to do a sum, how does it know the calculator didn't make a mistake? The paper suggests we need better ways for the AI to check the work of these external tools so it doesn't get fooled.
5. The Human-AI Dance
The paper argues that the future isn't "AI vs. Human." It's AI + Human.
- The Fix: Think of the AI as a super-fast research assistant. It can read thousands of papers in a second, check 1,000 possible proofs, and organize the data. The human is the Captain, providing the intuition, the "gut feeling," and the final judgment on whether an idea is actually brilliant or just a hallucination.
- Real-world example: In late 2025, an AI agent named Aristotle helped solve a problem from the Erdős list. It generated a proof, but it turned out the problem it solved was a slightly "weakened" version of the original. The human mathematicians had to step in to realize the original problem was still open. This shows that while AI is powerful, it still needs a human to keep it honest.
The Bottom Line
The paper is a call to action. It says: "We have built amazing Solvers. They can win any math competition. But to push the boundaries of human knowledge, we need to build Research Agents."
These agents won't just follow the rules; they will help us write new rules. But the paper is clear: we aren't there yet. The AI is currently a brilliant apprentice, not a master. It needs better data, better maps, better tools, and most importantly, a human partner to guide it through the dark, uncharted territories of mathematics.
The journey from "Solving" to "Discovering" is the next big leap, and it's going to take a team effort to get there.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.