← Latest papers
🤖 AI

OEIS Open: How many conjectures can language models turn into theorems?

This paper introduces OEIS Open, a secure benchmark of 492 formalized mathematical conjectures from the OEIS, demonstrating that language models equipped with minimal tools can autonomously resolve approximately 30% to 44% of these open problems at a modest cost, though access to extensive literature and sophisticated agent loops did not significantly improve performance.

Original authors: Tom Adamczewski

Published 2026-08-13
📖 4 min read☕ Coffee break read

Original authors: Tom Adamczewski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a vast, digital library of numbers called the OEIS (the Online Encyclopedia of Integer Sequences). It's like a giant catalog where mathematicians list patterns they've found in numbers, from the simple (like 1, 2, 4, 8...) to the bizarre and mysterious. Often, after listing a pattern, someone will write down a guess about how it works forever. These guesses are called "conjectures." For a long time, proving these guesses was a job only for human geniuses with chalkboards and endless patience. But recently, computers have started trying to solve these puzzles too. The big question is: Can an artificial intelligence (AI) actually figure out the truth on its own, or does it just need a human to hold its hand? This paper dives into that question by setting up a rigorous test where AI agents have to prove or disprove these number guesses without any unauthorized methods or human nudging.

The researchers behind this study, from a group called Epoch AI, created a new challenge called OEIS OPEN. Think of it as a giant math obstacle course. They took 492 open mathematical guesses from the OEIS and translated them into a strict computer language called Lean, which acts like a super-strict referee. In this game, an AI agent is dropped into a digital room with a few basic tools: a text editor, a command line (like a computer terminal), and a calculator. The AI's only goal is to write a proof that the computer referee will accept. If the AI can't prove the guess is true, it must prove it's false. The catch? The AI has a strict budget of $50 per guess to spend on computer time. If it runs out of money before solving it, it loses.

The results were surprisingly promising. The best AI models managed to solve 147 of the 492 guesses, scoring about 30% on the test. That means they turned nearly a third of these open mysteries into confirmed theorems (or disproofs) all by themselves. The most successful model, Claude Opus 4.8, solved 30% of them, while others like GPT-5.5 and Gemini 3.5 Flash also did well, solving 26% and 22% respectively. This is a big deal because the researchers used a very simple AI setup—just a basic loop that tries, checks, and tries again. It wasn't a fancy, super-complex robot with a massive team of human helpers. In fact, this simple approach actually did better than a much more complicated system called AlphaProof Nexus, which only solved 9% of the same problems.

The researchers also tested some "power-ups" to see if they would help the AI get smarter. They gave the AI access to a massive library of 476,000 math papers from the internet (arXiv), hoping it could learn from past human work. They also tried giving the AI a more complex "brain" that could delegate tasks to sub-agents and remember things over a long time. Surprisingly, neither of these upgrades helped. The AI didn't solve any more guesses with the library, and the complex brain didn't make it faster or more accurate. It seems that for this specific type of math problem, having a simple, focused toolset and a good budget is more important than having a massive library or a complicated personality.

One interesting finding was that the more money the researchers were willing to spend on a single guess, the more likely the AI was to solve it. The success rate went up in a steady, predictable way: for every ten times they increased the budget, the success rate jumped by about ten percentage points. This suggests that if they had given the AI a bigger budget, it could have solved even more. However, the paper is careful to note that these guesses, while mathematically valid, are mostly "unknowns" in the world of math. They aren't famous, world-changing problems like the Riemann Hypothesis; they are likely small, obscure puzzles that haven't received much attention from human mathematicians.

So, what does this all mean? The paper shows that AI is now capable of autonomously solving real, open mathematical research problems at a modest cost. It's not just solving puzzles with known answers anymore; it's pushing the frontier of what we know. But it also shows limits: the AI didn't get smarter just by reading more books, and it still struggles with the hardest, most obscure problems. The researchers conclude that while we aren't seeing a "qualitative jump" where AI suddenly becomes a genius mathematician overnight, we are seeing a steady, powerful tool that can chip away at the unknown, one number sequence at a time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →