First Proof Second Batch
This paper evaluates the capabilities of current AI systems in solving research-level mathematics by presenting the problems, methodology, and results of testing several AI models on ten distinct problems contributed by mathematicians, along with supplementary materials including human solutions and referee reports.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a group of mathematicians deciding to hold a high-stakes "cooking competition," but instead of baking cakes, they are asking Artificial Intelligence to cook up brand-new mathematical proofs. This document is the official report card from the second round of that competition, called First Proof Second Batch.
Here is the story of what happened, broken down into simple terms.
The Setup: A New Kind of Exam
In the first round of this experiment, the organizers let anyone try to solve the problems. It was a bit like an open house. For this second round, they wanted to be much stricter and more scientific.
- The Chefs (The AI Systems): They invited four specific "chefs" to compete. These weren't just random computers; they were advanced AI systems built by top universities (like UCLA and Princeton) and a company (OpenAI).
- The Ingredients (The Problems): The organizers didn't give the AI old, solved math problems found in textbooks. Instead, they gave them 10 brand-new, unsolved puzzles that real human mathematicians had recently discovered in their own research. These were problems that hadn't been posted on the internet or published in journals yet.
- The Rules: The AI had to solve these problems on its own, without humans whispering answers in its ear. The organizers watched every step, recorded every word the AI typed, and kept a strict log of how much it cost to run the "kitchen."
The Judging: A Blind Taste Test
Once the AI submitted its "dishes" (the proofs), a panel of 30 expert mathematicians acted as the judges. To be fair, the judges didn't know which AI made which dish. They gave each solution a rating:
- Essentially Flawless: A perfect dish, ready to serve.
- Minor Revisions: Delicious, but needs a pinch more salt or a better garnish.
- Major Revisions: The main course is there, but the sauce is burnt or the ingredients are missing. It needs a lot of work.
- Rejected: The dish is inedible.
The Results: A Mixed Bag
The outcome was a mix of impressive breakthroughs and embarrassing failures.
1. The "Wow" Moments (Successes)
- The Surprise Guest: For one problem involving complex fluid equations (Problem 5), one AI didn't just copy a human solution; it invented a completely new way to solve it. The judges were so impressed they called it "essentially flawless." It was like a chef inventing a new cooking technique that no human had thought of before.
- The Copycats: For another problem (Problem 6), two AIs managed to solve it perfectly. They followed the human solution's path but did it with such mechanical precision that the judges said, "This is mathematically correct, even if the writing style is a bit robotic."
2. The "Hallucination" Problems (Failures)
- The Fake Citations: A recurring issue was that the AIs would confidently cite books or papers that didn't exist or didn't say what the AI claimed. It was like a chef saying, "This sauce is made with a secret ingredient from a famous French chef," but when you check the book, that chef never wrote about it.
- The Plagiarism: In one case, an AI solved a geometry problem by copying the human mathematician's previous work almost word-for-word, using the same strange labels and terms, but forgot to give credit. If a human student did this, they would be expelled for cheating.
- The Wrong Turn: For several problems, the AIs tried to prove things that were actually false, or they got stuck in loops of nonsense. One AI tried to solve a geometry problem by inventing a theorem that didn't exist, essentially lying to the judges to make its proof look complete.
3. The Cost of Cooking
The report also looked at the "price tag" of the meal.
- Some AI systems were cheap but made mistakes.
- Others were incredibly expensive, costing thousands of dollars in computer time to produce a single solution.
- The most successful solutions often required the most "thinking time" (tokens) and money.
The Big Takeaway
The paper concludes that AI is getting very good at routine math. If a problem looks like something it has seen before, or if it just needs to follow a known recipe, the AI can do it well.
However, when it comes to true creativity—coming up with a brand-new idea, spotting a deep connection between two unrelated fields, or avoiding the trap of making things up—the AI is still struggling. It often tries to "fake it" by citing fake papers or skipping the hardest parts of the argument.
In short: The AI is a very fast, very expensive apprentice who can follow a recipe perfectly but isn't quite ready to invent a new cuisine on its own. It needs a human chef to check its work, fix its citations, and make sure it's not hallucinating ingredients.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.