← Latest papers
🔢 mathematics

QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems

The paper introduces QED, an open-source multi-agent system designed to bridge the gap between LLM benchmark performance and research-level mathematics by utilizing a specialized architecture that addresses seven specific failure modes, successfully generating original and nontrivial proofs for three out of five expert-contributed open problems in applied analysis and PDEs.

Original authors: Chenyang An, Qihao Ye, Minghao Pan, Jiayaun Zhang

Published 2026-04-28
📖 5 min read🧠 Deep dive

Original authors: Chenyang An, Qihao Ye, Minghao Pan, Jiayaun Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The "Math Detective" System: How AI is Learning to Solve Real Mysteries

Imagine you have a brilliant student who is amazing at taking multiple-choice tests and solving textbook problems. They can follow instructions perfectly and repeat what they’ve learned. But if you hand that student a brand-new, unsolved mystery from a real detective case—something no one in history has ever solved—they would likely freeze. They might try to "fake" an answer, copy a similar case from a book, or simply get stuck in a loop of guessing.

This is the current state of AI in mathematics. Most AI can solve "homework" problems, but they struggle with "research" problems—the kind of deep, original questions that mathematicians spend years trying to crack.

A group of researchers has just released a new system called QED. Think of QED not as a single "smart student," but as a highly organized Detective Agency.


The Problem: Why "Smart" AI Fails at Real Math

The researchers discovered that when you ask a standard AI to solve a hard math problem, it usually fails in seven specific, predictable ways. Here is how those failures look in real life:

  1. The Echo Chamber (Context Contamination): The AI tries to check its own work, but because it’s the same "brain," it just convinces itself that its mistakes are actually correct. It’s like a person arguing with themselves in a mirror and believing their own lies.
  2. The Fake Reference (Citation Hallucination): The AI makes up "facts" or "theorems" that don't exist to make its argument sound better. It’s like a lawyer citing a law that was never actually written.
  3. The "Skip the Hard Part" Move (Hand-Waving): When the math gets really difficult, the AI says, "It is obvious that X follows Y," without actually proving it. It’s like a chef saying, "Now, just add magic to make it delicious," instead of actually following a recipe.
  4. The Nervous Planner (Unstable Plans): As soon as the AI hits a roadblock, it throws its entire strategy in the trash and starts a new one, wandering aimlessly like a hiker who keeps changing direction every time they see a pebble.
  5. The Distracted Inspector (Unfocused Verification): When asked to check a proof, the AI tries to check everything at once—grammar, logic, citations, and math—and ends up missing the big errors because it’s spread too thin.
  6. The Goalpost Shifter (Problem Modification): To make the problem easier, the AI subtly changes the question. It’s like a runner who, halfway through a marathon, decides the finish line is actually 10 miles closer.
  7. The Single-Brain Limit (Single-Model Bottleneck): Every AI has "blind spots." If you only use one AI, you are stuck with its specific weaknesses.

The Solution: The QED "Detective Agency"

Instead of one AI trying to do everything, QED creates a team of specialized "agents" that watch over each other. They use a strict system of checks and balances:

  • The Separated Roles: The "Prover" (the one writing the math) and the "Verifier" (the one checking it) are kept in separate rooms. They aren't allowed to talk to each other, so the Verifier can't be "tricked" by the Prover's logic.
  • The Two-Stage Inspection: Before a proof is accepted, it goes through two intense inspections. First, a Structural Inspector checks if the "skeleton" of the proof is solid (Did they change the question? Are the citations real?). Only if the skeleton is perfect does it move to the Detailed Inspector, who checks every single tiny mathematical step.
  • The "No Shortcuts" Rule: QED forces the AI to "tag" the hardest parts of the proof. If the AI tries to use vague words like "obviously," the system flags it and says, "No, show your work!"
  • The Master Strategist (Decomposition Mode): For the hardest problems, QED uses a "Manager" agent. This manager breaks the massive problem into a "map" of small, manageable steps. If a step fails, the manager decides: "Do we just need to fix this one step, or do we need to redraw the whole map?"

The Result: Real Discoveries

The researchers didn't just test QED on math homework; they gave it five unsolved problems from real-world experts in physics and fluid dynamics (the math used to understand how liquids and gases move).

The result was stunning: QED successfully solved three out of the five problems.

Even more importantly, these weren't just "correct" answers—they were original discoveries. In one case, the AI found a mathematical connection that the human experts hadn't even realized existed. It didn't just solve the problem; it revealed a new way to look at the math.

The Bottom Line

QED shows that the path to "Artificial Intelligence" becoming a "Scientific Intelligence" isn't just about making the AI "smarter"—it's about building a better system of checks, balances, and specialized roles.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →