MathAtlas: A Benchmark for Autoformalization in the Wild
This paper introduces MathAtlas, the first large-scale autoformalization benchmark comprising approximately 52,000 graduate-level mathematical concepts from 103 textbooks, which reveals the extreme difficulty of current models in handling complex, dependency-rich research mathematics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a brilliant translator who can turn human stories into a strict, computer-readable language called "Lean." For a long time, this translator has only been tested on simple stories—like high school math problems or basic puzzles. The computer could easily translate these because the rules were simple, and the computer had already seen similar rules before.
But what happens when you ask this translator to translate a complex, graduate-level research paper? That is the problem MathAtlas tackles.
Here is a simple breakdown of what the paper does, using some everyday analogies:
1. The Problem: The "Library" is Too Big
Imagine you are trying to build a house (a formal proof) based on a blueprint (a math textbook).
- Old Benchmarks: Previous tests only asked the translator to build a small shed. The materials were all right there in the toolbox, and the instructions were short.
- The Real World: Graduate math is like trying to build a skyscraper. To build the top floor, you need the 50th floor. To build the 50th floor, you need the 49th, and so on, all the way down to the foundation.
- The Issue: In advanced math, the "foundation" (definitions and theories) often hasn't been built in the computer's language yet. If the translator doesn't know the definition of a "Lie Algebra" (a complex concept), it can't translate the theorem that uses it.
2. The Solution: MathAtlas (The "Wild" Map)
The authors created MathAtlas, a massive new test set.
- The Scale: They scraped 103 graduate-level math textbooks. It's like taking a library of 52,000 pages of advanced math and turning it into a digital map.
- The "Dependency Graph": This is the paper's superpower. Imagine a "Choose Your Own Adventure" book where every page has arrows pointing to other pages you need to read first. MathAtlas draws these arrows. It shows that to understand Proposition 15, you first need to understand Dedekind Rings, which in turn needs Prime Ideals.
- Why it matters: It forces the AI to not just translate one sentence, but to figure out: "Do I have the tools to build this? If not, can I find or build the tools first?"
3. The Results: The AI is Stuck in the Mud
The authors tested the smartest AI models available on MathAtlas, and the results were humbling.
- The Score: Even the best AI got less than 10% of the theorem statements correct. For definitions, it was around 16%.
- The "Depth" Trap: The deeper the "dependency tree" (the more steps you have to go back to find the definitions), the worse the AI performed.
- Analogy: If the AI has to look up 10 previous concepts to translate one sentence, it gets confused and gives up. On the hardest problems (where the dependency tree was deepest), the AI got only 2.6% correct.
- The "Mathlib" Effect: The AI did slightly better when the concept it needed to translate was already in a famous library called "Mathlib" (which is like a pre-built toolbox the AI has likely studied). If the concept was new and not in that library, the AI struggled much more.
4. The "Trust" Test (MA-Align)
The paper also created a new test called MA-Align.
- The Problem: Sometimes an AI translates a sentence into code that looks perfect and runs without errors, but it actually means something completely different than the original math. It's like translating "The cat sat on the mat" into a computer program that says "The dog ate the pizza." The program works, but it's wrong.
- The Test: They asked AI judges to check if the translation was "faithful" (true to the meaning). They found that even the best AI judges were often fooled, especially with complex graduate-level definitions.
The Bottom Line
The paper concludes that while AI is getting good at simple math, it is currently terrible at translating complex, real-world graduate mathematics. The main bottleneck isn't just translating words; it's understanding the massive web of connections between ideas and knowing how to build the necessary "tools" (definitions) before building the "house" (theorems).
MathAtlas is now open for everyone to use as a challenge course to help AI learn how to handle these deep, complex dependencies.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.