GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
The paper introduces GAUGE, a novel benchmark that evaluates AI-built financial models against observed analyst practices rather than a single "golden" answer, revealing that while current agents excel at mechanical model construction, they still significantly lag behind human professionals in valuation judgment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a judge on a cooking show. Usually, the judges have a single "perfect" recipe card. If a contestant's dish tastes even slightly different from that card, they lose points. But what if there are actually ten different ways to make a delicious lasagna? What if the "perfect" recipe is just one person's opinion, and the contestant's version is equally tasty but uses different spices? This is the problem scientists face when testing AI on complex tasks like financial modeling. Financial modeling is like building a detailed prediction machine for a company's future; it mixes hard facts (like past sales) with educated guesses (like how fast the company will grow). For decades, we've tested AI by comparing its guesses to a single expert's answer. But if experts themselves disagree on the best guess, punishing the AI for having a different, yet reasonable, answer doesn't seem fair. This paper asks: How do we grade AI when there isn't just one "right" answer?
The researchers behind this study, titled GAUGE, decided to stop pretending there is only one perfect financial model. They built a new testing ground to see if AI agents can build these complex financial "prediction machines" and, more importantly, if they can make the judgments inside them that human experts make.
First, they tested the old way of grading. They took 108 pairs of financial models built by different human experts for the same companies. When they graded one expert's model against another's using the old "single perfect answer" rule, the results were shocking. The median score was only 0.33 out of 1.0. In fact, 92.6% of the pairs scored below 0.70. This means that even professional humans, who are all experts, often disagree so much that if you treated one as the "truth," you would fail the other. It turns out that in finance, being "different" doesn't mean being "wrong."
To fix this, the team created GAUGE (Grading Agent-Built Financial Models Without a Golden Answer). Instead of a single recipe card, they created a "safety envelope" based on what real experts actually do. Imagine a target where the bullseye isn't a single dot, but a wide ring. If an AI's guess lands anywhere inside that ring of "reasonable expert behavior," it gets credit. The system checks two things:
- The Mechanics: Did the math work? Did the spreadsheet balance? (Like checking if the oven was actually turned on).
- The Judgment: Did the AI's guesses fall within the range of what real experts consider defensible? (Like checking if the seasoning is within the range of a good chef's taste).
They tested 24 different AI agents on this new system, along with a group of 55 humans ranging from finance students to senior analysts. Here is what they found:
- AI is getting good at the math, but bad at the gut feeling. The AI agents were excellent at the mechanical parts. The best AI passed 93% of the mechanical checks (making sure the spreadsheet formulas were correct). However, when it came to the judgment parts (making the actual predictions), the best AI only passed 78%.
- The "Gap" is real. On average, the AI agents were 26 points better at building the model than at making the valuation judgments. They can build the car, but they aren't great at driving it yet.
- Humans still win. The best AI scored 53.4. This was better than the average finance student (43.2) but worse than every single senior analyst (who averaged 88.3) and most junior analysts. The AI is smart enough to do the homework, but not quite smart enough to be the boss.
The study suggests that while AI is becoming very capable of constructing complex financial models, the hardest part—making the nuanced, defensible judgments that experts use to value a company—remains a significant challenge. The authors didn't just find a score; they proved that the old way of grading (looking for one perfect answer) was unfair to both humans and machines, and they offered a new way to measure progress that respects the messy reality of expert disagreement.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.