RMA: an Agentic System for Research-Level Mathematical Problems
The paper introduces Research Math Agents (RMA), an agentic framework that outperforms strong baselines like GPT-5.2R on research-level mathematical problems by utilizing a multi-role, iterative workflow to decompose tasks into specialized modules for literature grounding, proof generation, and verification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a brand-new, unsolved mystery in the world of mathematics. It's not like a puzzle from a high school textbook where the answer is hidden in the back of the book, and you just need to follow a known recipe. Instead, it's like trying to write a new chapter for a dictionary that hasn't been written yet. You have to invent the rules, find the right words, and prove your ideas are true, all while making sure you haven't accidentally copied someone else's work.
This paper introduces RMA (Research Math Agents), a new team of AI "robots" designed specifically to tackle these tough, unsolved math mysteries.
Here is how RMA works, explained through a simple analogy:
The Problem: The "Lone Genius" vs. The "Research Team"
Previous AI systems were like lone geniuses. They were great at solving competition math problems (like the International Math Olympiad) because those problems have known answers and clear paths. But when faced with real, open-ended research math, they often got stuck, made up facts (hallucinations), or gave up.
RMA changes the game by acting like a professional research team in a university lab. Instead of one robot trying to do everything alone, RMA breaks the work down into a structured team effort.
The RMA Team: A Division of Labor
Think of RMA as a construction crew building a skyscraper (the proof) from scratch. They don't just guess; they follow a strict workflow with three main roles:
The Initializer (The Architect):
- Job: This agent looks at the messy, confusing problem statement and turns it into a clear blueprint. It breaks the big problem into smaller, manageable goals.
- Analogy: Like an architect who takes a client's vague idea ("I want a building that floats") and draws the first set of blueprints, defining exactly what needs to be built.
The Proposers (The Builders & Engineers):
- Job: These agents take the blueprint and try to build the proof. If they hit a wall (a logical gap), they don't just give up. They go back to the "library," find a new tool or a similar building design (a theorem from a textbook), and try a different construction method.
- Analogy: Like a team of engineers trying different materials and designs. If a beam breaks, they don't just say "it's impossible"; they check the library for a stronger steel alloy and try again.
The Verifiers (The Inspectors):
- Job: These agents act as strict building inspectors. They don't build anything; they only check the work. They look for cracks in the logic, missing steps, or fake materials. If they find a mistake, they send the builders back to fix it.
- Analogy: Like a safety inspector who walks through the building, tapping walls and checking blueprints. If they find a flaw, they issue a "fix-it" ticket, and the builders must correct it before moving on.
The "Shared Notebook" (Structured Memory)
A key part of RMA is that everyone shares a digital notebook (a shared memory).
- The Architect writes the blueprint.
- The Builders add their construction notes and new ideas.
- The Inspectors write their critique notes.
- Crucially: No one erases what came before. They just add new pages. This means the team remembers every mistake they made and every tool they found, so they don't repeat errors.
The "Fair Play" Rules
To make sure the AI isn't cheating by looking up the answer, RMA has a Fair Comparison Module.
- It acts like a strict referee who locks the doors to any website that might have the solution.
- It ensures the AI only looks at old textbooks and papers that were published before the problem was even created, forcing the AI to actually think and discover rather than just copy-paste.
The Results: Did It Work?
The researchers tested RMA on the "First Proof" benchmark, which consists of 10 real, unsolved math problems contributed by famous mathematicians.
- The Competition: They compared RMA against other top AI systems, including "GPT-5.2R" (from OpenAI) and "Aletheia" (from Google DeepMind).
- The Outcome:
- Aletheia couldn't solve any of the problems.
- GPT-5.2R solved 3 out of 10, but sometimes made small logical errors or gave weaker answers.
- RMA solved 8 out of 10 problems.
- More importantly, when human mathematicians reviewed the proofs, they said RMA's solutions were more logical, clearer, and more trustworthy than the others.
The Bottom Line
The paper claims that solving hard math isn't about having a smarter "brain" (a single powerful AI model); it's about having a better process. By breaking the work into specialized roles (Architect, Builder, Inspector), giving them a shared notebook, and forcing them to check their work repeatedly, RMA can tackle problems that even the smartest single AI cannot.
It's the difference between asking one person to build a house alone in the dark versus hiring a team with a blueprint, a library, and a safety inspector.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.