Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics
The paper introduces "Formal Conjectures," a dynamic, open-source benchmark of 2,615 Lean 4 formalized mathematical problems sourced from active research, designed to evaluate and drive the capabilities of automated reasoning systems in discovering new proofs while ensuring data integrity through community collaboration and AI-audited verification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to do advanced mathematics. In the past, you might have given the robot a stack of old math homework problems. But there's a big problem: the robot might have just memorized the answers from the internet instead of actually learning how to think. It's like a student who memorized the answer key to a test but doesn't understand the math.
This paper introduces a new, smarter way to test these robots. They call it Formal Conjectures.
Here is how it works, broken down into simple ideas:
1. The "Fresh Meat" Test (Zero Contamination)
Most math tests for AI are "stale." The answers are already online, so the AI might just be copying them.
- The Analogy: Imagine a cooking competition where the judges give the chefs a recipe they've never seen before, written in a secret code. If the chef can cook the dish, they actually know how to cook. If they can't, they are just guessing.
- The Paper's Solution: The authors created a library of 1,029 unsolved math problems (conjectures). These are problems that real human mathematicians are currently trying to solve. Because no one has solved them yet, the AI cannot have memorized the answers. If the AI solves one, it's a genuine discovery, not a copy-paste job.
2. The "Strict Judge" (Lean 4)
In normal math, you can write a proof that looks good but has a tiny logical error. Humans might miss it.
- The Analogy: Think of a video game where you have to build a bridge. If you use a weak brick, the bridge collapses. In this paper, the "bridge" is a mathematical proof. The authors use a special computer language called Lean 4 as the judge.
- The Paper's Solution: Lean 4 is like a super-strict referee. It doesn't care if your proof looks pretty; it checks every single logical step. If there is even one tiny mistake, the referee says, "Nope, that's wrong." This ensures that when the AI says it solved a problem, it actually did.
3. The "Living Library" (An Evolving Benchmark)
Usually, once a test is made, it stays the same forever. But AI gets smarter every day, so old tests become too easy.
- The Analogy: Imagine a video game that updates every week. As players get better, the game adds harder levels and fixes bugs in the map.
- The Paper's Solution: This benchmark is a "living" project.
- It grows: They keep adding new problems from real research papers.
- It fixes itself: Sometimes, the math problems themselves are written vaguely. When the AI tries to solve them, it might fail because the problem was unclear. This failure helps human mathematicians realize, "Oh, we wrote that problem wrong!" They then fix the problem statement.
- It has two tracks:
- The Discovery Track: Trying to solve the unsolved problems (the "fresh meat").
- The Translation Track: Taking problems that humans have already solved and teaching the AI how to write them in the strict computer language (Lean 4).
4. The "Safety Net" (Avoiding Cheating)
The authors are worried about "data leakage" (the AI cheating by seeing the answers in its training data).
- The Analogy: To stop students from cheating, teachers sometimes lock the answer key in a safe and only open it after the test.
- The Paper's Solution: They created a "frozen" version of the test. This is a snapshot of 100 problems that is locked in time. Even if the main library changes later, this specific snapshot stays the same so scientists can compare different AI models fairly without worrying about the test changing underneath them.
5. Why This Matters
The paper shows that this system is already working.
- Real Results: They mention that using this system, an AI (called Aristotle) helped a human mathematician solve a famous problem that had been open for a long time (Erdős Problem 124).
- The Signal: The benchmark provides a clear "climbable signal." It shows exactly how far AI has come and how much further it needs to go. If an AI solves 10% of the unsolved problems, that's a huge deal. If it solves 0%, it's not ready yet.
Summary
Formal Conjectures is a new, ever-updating playground for AI mathematicians. It uses a strict computer referee (Lean 4) to check work, focuses on problems that haven't been solved yet to prevent cheating, and acts as a collaborative tool where humans and AI help each other clarify difficult math questions. It's not just a test; it's a tool for making real mathematical discoveries.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.