Mathematical Scientific Discovery Using Large Language Models: A Systematic Literature Review
This systematic literature review, adhering to PRISMA 2020 guidelines, analyzes 32 studies on using large language models for mathematical discovery to identify seven technical paradigms, highlight the superior efficacy of evolutionary search combined with formal verification, and provide a taxonomy to guide future research in the field.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where math isn't just about solving puzzles someone else wrote down, but about inventing the puzzles themselves. For centuries, the most exciting part of mathematics—the "Eureka!" moment of discovering a brand-new truth, a strange new pattern, or a rule that no one has ever seen before—was thought to be a superpower unique to humans. It required a special kind of intuition, a spark of creativity that computers just couldn't have. But recently, a new kind of artificial intelligence called a "Large Language Model" (or LLM) has started to change the game. Think of an LLM not as a calculator, but as a super-reading machine that has swallowed almost every math book, code manual, and scientific paper ever written. It doesn't just memorize them; it learns the "grammar" of math and code so well that it can start writing its own sentences. The big question researchers are asking is: Can this machine-reading super-brain do more than just answer questions? Can it actually discover new math that humans haven't found yet?
This is exactly what a team of researchers from Clemson University and Colorado College set out to investigate in a massive new study. They didn't just look at one cool experiment; they went on a digital treasure hunt, sifting through hundreds of research papers published up to the end of 2025 to find the ones where AI was actually creating new mathematical knowledge. They used a strict, scientific checklist (called PRISMA) to make sure they didn't miss anything important and to filter out studies that were just about solving old problems. After the dust settled, they found 32 studies that showed something truly remarkable: AI is starting to become a partner in mathematical discovery, not just a calculator.
The researchers found that the most successful AI "mathematicians" aren't just guessing randomly. Instead, they use a few clever tricks. One of the most powerful methods is like a digital version of evolution. Imagine a computer creating thousands of tiny, simple math programs. It tests them, keeps the ones that work best, and then uses the AI to "mutate" them—making small, creative changes to see if they get even better. Over time, this process evolves solutions that are so good they beat records humans have held for decades. For example, the study highlights how these systems recently found a faster way to multiply matrices (a specific type of math grid) that improved upon a famous algorithm from 1969, something no human had managed to do in over 50 years.
Another popular method is like a video game where the AI plays against itself. One part of the AI tries to create a difficult math problem (a "conjecture"), and another part tries to solve it. As the solver gets better, the problem-maker has to get smarter, creating a loop where they both level up together. This has helped AI generate thousands of new, verified math theorems and lemmas (small stepping-stone proofs) that are now being added to digital libraries. However, the paper is very clear that these AI systems aren't magic wands yet. They are still very hungry for computer power, often needing millions of tries to find one good answer. Also, while they are great at finding patterns, they sometimes produce answers that are technically correct but boring or useless to human mathematicians.
The study also points out that for these discoveries to be trusted, they need a "referee." In the world of math, you can't just say "I think this is true." You have to prove it. The most popular referee in these AI experiments is a tool called "Lean," a digital proof-checker that acts like a strict grammar police for math. If the AI says it found a new theorem, Lean checks every single step to make sure it's 100% logical. The researchers found that the most successful projects were the ones where the AI and the referee worked together in a tight loop, with the referee giving immediate feedback to the AI so it could fix its mistakes on the fly.
So, what's the bottom line? The paper suggests that we are standing on the edge of a new era. AI isn't replacing mathematicians, but it is becoming a powerful "idea sandbox." It can churn out millions of possibilities, spot patterns in massive amounts of data that humans would miss, and even solve problems that have been stuck for half a century. But the final step—turning a machine-generated code snippet into a beautiful, understandable mathematical insight—still needs a human touch. The future looks like a team sport: the AI does the heavy lifting of searching and testing, and the human mathematician provides the intuition to decide what's actually interesting. It's a partnership where the machine brings the speed and the scale, and the human brings the soul.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.