← Latest papers
📈 economics

Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research

This paper proposes a verification-first protocol for using large language models in economic theory research, demonstrating through a single worked example that external adversarial verification and structured multi-agent checks are essential for ensuring rigor, as model fluency alone does not guarantee mathematical correctness.

Original authors: Moran Koren

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Moran Koren

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a complex machine, like a clockwork engine, but you have a very talented, confident robot assistant who can write the blueprints and do the math for you. The problem is, this robot is a "smooth talker." It can write a perfect-looking blueprint for a machine that doesn't actually work. It might say, "Here is the gear that balances the weight!" when, in reality, that gear is just a drawing.

This paper is a guide on how to use that robot assistant without getting fooled by its confidence.

The Problem: The "Smooth Talker" Trap

In the world of economics theory, researchers usually start with a blank page and have to do all the hard math themselves. Now, with AI (Large Language Models), the robot can do the math and write the proof. But the robot has a flaw: it is too confident. It will prove a false theorem just as fluently as a true one. It's like a student who memorizes the answer key but doesn't understand the math; if you ask a hard question, they might just guess the right-sounding answer.

The author, Moran Koren, asks: How do we use this robot to do real economic theory without trusting it blindly?

The Solution: The "Verification Toolbox"

The paper doesn't try to make the robot smarter. Instead, it proposes three different ways to check the robot's work. Think of these as three different ways to hire a team to build your clockwork engine.

Method 1: The Solo Artist (The Single Pass)

You ask the robot to do the whole job in one go. It writes the proof, checks its own work, and hands you the final paper.

  • The Analogy: It's like asking a single, very confident architect to design a bridge. They check their own math, and if it looks clean, they sign it.
  • The Result: The output looks beautiful and finished. But because there was no one else to check it, it might have hidden cracks. In the paper, this method produced the "prettiest" paper, but it was the least verified.

Method 2: The Debaters (The Adversarial Pair)

You hire two robots. One is the Builder (who tries to prove the math), and the other is the Hater (whose only job is to try to break the Builder's work). A human acts as the referee.

  • The Analogy: Imagine a courtroom. The Builder presents a case. The Hater is a lawyer whose job is to find a single flaw to tear it apart. If the Hater can't find a flaw, the human referee checks the evidence.
  • The Result: This method caught three major errors that the "Solo Artist" missed. For example, the Builder claimed a machine was balanced, but the Hater proved it would collapse under specific conditions. The Hater forced the team to admit, "Okay, we were wrong about that part."

Method 3: The Research Lab (The Multi-Agent Project)

You set up a whole team of robots with specific roles: one does the literature, one does the math, one writes code, and one is a strict Reviewer who acts as a gatekeeper. Nothing gets published unless the Reviewer signs off.

  • The Analogy: This is like a rigorous scientific lab. Every step is documented. If a robot makes a guess, it has to put a big red flag on it saying, "I haven't proven this yet." The Reviewer stops the project if a step isn't solid.
  • The Result: This produced the messiest-looking document. It was full of red flags, notes saying "we don't know this yet," and rejected ideas. But because it was so honest about what it didn't know, it was actually the most trustworthy.

The Test: The "Grade Inflation" Puzzle

To test these three methods, the author gave them a very hard, unsolved puzzle about grade inflation.

  • The Puzzle: Imagine a university where students take different courses. Some courses are easy, some are hard. How do you create a fair system to measure a student's true ability, regardless of which classes they took?
  • The Goal: The robots had to design a system (a "mechanism") that would stop students and teachers from cheating the system to get better grades.

What Happened?

The author ran all three methods on this same puzzle. Here is what they found:

  1. Convergent Discovery (The "Aha!" Moment): Two of the methods, which never talked to each other, independently discovered the exact same mathematical formula. It was like two different detectives solving a crime and finding the same hidden clue. This gave the author confidence that the clue was real and not just a lucky guess.
  2. The Hater is Essential: The "Debater" method (Method 2) caught three specific lies that the "Solo Artist" (Method 1) would have published as facts. The Hater saved the day by proving the machine wouldn't work as claimed.
  3. Polish is Not Proof: The most beautiful, finished-looking paper (Method 1) was the most dangerous because it had no safety checks. The messiest paper (Method 3), covered in "I'm not sure" flags, was the safest because it told the truth about its limitations.

The Big Takeaway

The paper concludes that how you check the work matters more than how smart the robot is.

If you want to use AI for serious economic theory, you shouldn't just ask it to "write a proof." You need to set up a system where:

  • Someone tries to break the proof.
  • Someone keeps a list of what hasn't been proven yet.
  • You run the same problem twice to see if you get the same answer.

The paper doesn't claim to have solved the grade inflation problem perfectly. Instead, it claims to have built a toolkit (a "Theorist Toolbox") that shows researchers how to use AI safely. The real result isn't the math about grades; it's the protocol for how to trust a machine that can lie so fluently.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →