MathDuels: Evaluating LLMs as Problem Posers and Solvers
The paper introduces MathDuels, a dynamic self-play benchmark where frontier language models simultaneously act as adversarial problem authors and solvers, utilizing a Rasch model to reveal decoupled capabilities and prevent evaluation saturation through a co-evolving difficulty curve.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out who the best chess player in the world is.
For a long time, we've done this by giving them a fixed list of 100 puzzles. If they solve 99 of them, they get a high score. But here's the problem: the puzzles are getting too easy. The top players solve them all perfectly, so we can't tell who is actually better. It's like trying to measure the height of giants using a ruler that only goes up to 6 feet; they all just hit the top and look the same.
This paper introduces a new way to test AI math brains called MathDuels. Instead of a static test, it turns math into a living, breathing arena where the AI models play two roles at once: The Architect and The Solver.
Here is how it works, using a simple analogy:
1. The Arena: A "Math Duel"
Imagine a tournament where every participant has to do two things:
- Role A: The Architect (Problem Creator): They must design a tricky math puzzle for everyone else to solve.
- Role B: The Solver (Problem Solver): They must try to solve the puzzles created by everyone else.
This is inspired by a real historical event from 1535. Two mathematicians, Tartaglia and Fior, didn't just take a test; they challenged each other. They wrote 30 problems for the other to solve. The winner wasn't just the one who could solve the best; it was the one who could create problems the other couldn't crack.
2. The Twist: "Self-Play"
In this AI tournament, the models play against each other in a giant round-robin.
- Model A writes 30 hard problems.
- Model B, C, D... try to solve Model A's problems.
- Then, Model B writes 30 problems, and everyone else tries to solve them.
The system uses a special "referee" (an independent AI) to check if the problems are fair and if the answers are actually correct. If a problem is broken or confusing, it gets thrown out.
3. The Scoreboard: Two Different Skills
The most surprising discovery of this paper is that being good at solving math doesn't mean you are good at making math.
Think of it like a video game:
- Some players are Speedrunners: They can finish any level incredibly fast (Great Solvers).
- Some players are Level Designers: They can build levels so tricky that even the speedrunners get stuck (Great Architects).
In the past, we only measured the Speedrunners. MathDuels measures both.
- The Result: One AI model (Gemini-3.1-Pro-high) wasn't the fastest solver, but it was the best architect. It built puzzles so clever that even the fastest solvers got stuck. Because of this, it ranked #1 overall.
- Another model (GPT-5.4-high) was the fastest solver but a mediocre architect. It ranked #2.
4. Why This Matters: The "Moving Target"
In old tests, once the AI gets too smart, the test becomes useless (it hits a "ceiling").
In MathDuels, the difficulty grows with the players.
- When a super-smart new AI enters the arena, it doesn't just solve the old puzzles; it creates new, harder puzzles that break the previous champions.
- It's like a video game where every time you beat the final boss, the boss upgrades itself to be even harder for the next round. The test never stops being challenging.
5. The "Trap" of Easy Problems
The researchers found that some models try to cheat by writing "trick" questions that look hard but are actually broken or unsolvable. The system has a safety net (the Verifier) that catches these. If a model writes a problem that is mathematically nonsense, it gets penalized. This forces the models to be honest and rigorous, not just tricky.
Summary
MathDuels is a new way to rank AI math skills by making them teach and test each other.
- It realizes that creating a hard problem is a different skill than solving one.
- It prevents the test from becoming "too easy" because the AI models keep inventing harder challenges for each other.
- It gives us a leaderboard that actually changes and evolves, showing us who is truly the "math genius" of the group, not just the one who memorized the answers.
It's the difference between a student who can ace a multiple-choice test and a professor who can design a test that stumps the entire class. MathDuels finds the professors.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.