← Latest papers
🤖 AI

InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

This paper introduces InvestPhilBench, a multi-layer dynamic benchmark designed to evaluate large language models' ability to reconstruct and apply expert investment decision frameworks, revealing that while composite automated scores often reward fluent prose, specialized procedural metrics expose significant reasoning deficits even in frontier models.

Original authors: Mingguang Chen, Bo Qu

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Mingguang Chen, Bo Qu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Can AI "Think" Like a Master Investor?

Imagine you hire a very smart, well-read assistant to help you manage your money. You ask them to analyze a company using Warren Buffett's specific method.

  • The Old Way: You check if the assistant knows facts about Buffett. Do they know he likes "economic moats"? Do they know he hates airlines? If they say "Buffett likes moats," they get a passing grade.
  • The New Problem: The paper argues that knowing the facts isn't enough. A true expert follows a strict decision process (a recipe). For Buffett, the first step is always: "Is this company inside my Circle of Competence?" If the answer is "No" (e.g., it's a biotech company), the process stops immediately. The deal is dead. No further analysis happens.

InvestPhilBench is a new test designed to see if AI assistants can actually follow these strict, step-by-step recipes, or if they just sound smart while skipping the most important steps.


The "InvestPhilBench" Test: An 8-Layer Obstacle Course

The researchers built a test with 8 levels of difficulty, like a video game with increasingly hard stages:

  1. Levels 1–3 (The Fact Check): Can the AI remember who said what? (e.g., "Who invented the 'Margin of Safety' concept?")
    • Result: The AI is great here. It's like a student who memorized the textbook.
  2. Levels 4–5 (The Recipe Reconstruction): Can the AI rebuild the investor's decision checklist in the correct order?
    • Result: This is where things get tricky. The AI starts to stumble.
  3. Levels 6–8 (The Real-World Simulation): Can the AI take a brand-new, weird investment scenario and apply the investor's specific rules to it?
    • Result: This is the hardest part. The AI often fails to apply the logic correctly, even if it knows the rules.

The "Magic" vs. The "Mechanics"

The paper discovered a surprising trick that the AI uses to cheat.

  • The "Fluent Prose" Trap: When the AI answers, it writes beautifully. It sounds confident and professional. If you just read the final paragraph, you might think it did a perfect job.
  • The "Gate" Reality: The researchers built a special tool (called BASP) to grade the answer. They found that while the AI gets high scores for "sounding good," it often misses the Kill Criteria.

The "Kill Criterion" Analogy:
Imagine a bouncer at a club.

  • The AI's mistake: The bouncer lets a person in because they are wearing a cool jacket (the "fluent prose").
  • The Reality: The bouncer should have checked the ID first. If the ID says "Under 21," the person gets kicked out immediately, no matter how cool the jacket is.
  • The Paper's Finding: The AI often skips the ID check (the "Kill Criterion") and lets the bad investment in, just because the explanation sounded nice.

The "Composite" vs. The "Gate" Score

The paper introduces two ways to grade the AI:

  1. The Composite Score (The "Fluency" Grade): This is like a teacher giving a student an 'A' because their essay was well-written, even if the math inside was wrong. The top AI models get very high scores here (around 90%).
  2. The Gate Reconstruction Accuracy (The "Logic" Grade): This is like a math teacher checking every single step of the calculation. When the researchers used this stricter test, the top AI models dropped significantly (down to around 57–77%).

The Takeaway: The AI is getting better at talking like an investor, but it is still bad at thinking like one.

The "Recipe" vs. The "Cookbook"

The researchers tested three ways to help the AI:

  1. Closed Book: The AI has to rely on its memory. (It struggles).
  2. The "Oracle" (The Full Cookbook): They gave the AI the entire rulebook for the investor to read while answering.
    • Surprise: Giving the AI the whole book actually made it worse in some cases. It got confused by too much information (like a chef trying to read the whole encyclopedia while cooking a steak).
  3. Targeted Retrieval (The Cheat Sheet): They gave the AI just the top 3 relevant rules.
    • Result: This worked best. It's like giving a chef a specific recipe card instead of the whole library.

The "Failure Modes" (How the AI Breaks)

The paper identified six specific ways the AI fails, which they named with catchy titles:

  • The "Hallucinated Gate": The AI invents a step in the process that doesn't exist (e.g., "Step 4: Check the moon phase" when the investor never said that).
  • The "Time Traveler": The AI mixes up dates. It might say Buffett hated airlines in 2020, when he actually bought them in 2016 and sold them in 2020. It flattens history into one confusing story.
  • The "Voice Changer": The AI sounds like a generic financial expert instead of the specific person. It might use Ray Dalio's words to explain George Soros's strategy.
  • The "Kill Criterion Omission": The most common failure. The AI lists the rules but forgets the "Stop" signs. It keeps analyzing a bad deal even after the first rule said "Reject."

The Bottom Line

InvestPhilBench is a new tool that proves current AI models are excellent at reciting investment philosophy but are still struggling to apply it.

  • What works: The AI can tell you what an investor believes.
  • What fails: The AI often cannot follow the strict, sequential logic of how that investor makes a decision, especially when it involves saying "No" to a deal.

The paper concludes that until AI can reliably follow these "kill criteria" and step-by-step logic, we cannot fully trust it to make complex investment decisions on its own. It's a tool for research, but it's not yet a replacement for the human logic of a master investor.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →