CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents
The paper introduces CLQT, a closed-loop, cost-aware, and strategy-consistent benchmark that shifts the evaluation of LLM portfolio-management agents from simple return-based ranking to a diagnostic framework capable of isolating specific reasoning failures and measuring durable competencies through a verifiable, multi-stage trading cycle.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a team of robot financial advisors to manage your money. In the past, people tested these robots by asking them simple questions like, "What is the price of Apple stock?" or by seeing who made the most money over a single quarter.
The authors of this paper argue that these old tests are flawed. They are like judging a chef solely by how many calories their dish has, without tasting the food or checking if they actually followed the recipe. A robot might make a lot of money just because the market happened to go up that day, not because it is actually smart. Or, it might cheat by "peeking" at tomorrow's news.
This paper introduces CLQT, a new way to test these AI agents. Think of CLQT not as a race to see who wins, but as a medical diagnostic tool for AI. Instead of giving a robot a single score (like "Gold Medal"), it gives the robot a detailed health report card with five specific vitals.
Here is how CLQT works, using simple analogies:
1. The "No Cheating" Rule (The Time Gate)
Imagine a test where students are allowed to look at the answer key before the exam starts. That's what many old tests did with AI; they accidentally let the AI see future data.
CLQT has a strict "Time Gate." It's like a security guard who checks the clock. The AI is only allowed to use information that existed before it made its decision. If it tries to peek at tomorrow's stock prices, the test immediately fails. This ensures the AI is actually thinking, not just remembering.
2. The Five-Part Health Check (The Scorecard)
Instead of asking "How much money did you make?", CLQT asks five deeper questions. It gives the AI a score from 0 to 100 on each:
- Coherence (The "Storyteller" Test): Does the AI's action match its own story?
- Analogy: Imagine a driver says, "I'm going to drive slowly because it's raining," but then they slam on the gas. CLQT checks if the AI's reasoning matches its actual trades. The paper found that many AIs tell a great story but do something completely different.
- Acuity (The "Focus" Test): Can the AI ignore the noise?
- Analogy: Imagine trying to hear a friend talk in a loud, chaotic stadium. Some AIs get distracted by every loud noise (random price jumps) and forget the important signal (real trends). CLQT measures if the AI can focus on the right things.
- Composure (The "Calmness" Test): Does the AI panic when things get shaky?
- Analogy: When the stock market drops suddenly, a calm investor stays steady. A panicked one sells everything in a frenzy. CLQT checks if the AI overreacts to small bumps in the road or stays disciplined.
- Discipline (The "Rule-Follower" Test): Does the AI stick to the plan?
- Analogy: If you tell a robot, "Only buy blue shirts," and it buys a red one, it lacks discipline. CLQT checks if the AI respects its own rules and budget limits without needing a human to yell at it.
- Reliability (The "Stamina" Test): Does the robot actually finish the job?
- Analogy: Some robots start a task, get confused, crash, or give up halfway through. CLQT counts how often the AI successfully completes its full cycle of thinking and acting without breaking down.
3. The "Two Modes" Experiment
The researchers tested the AIs in two different "personality" modes:
- Structured Mode: The AI acts like a strict committee. It has to follow a step-by-step checklist (like a pilot's pre-flight checklist).
- Autonomous Mode: The AI is given full freedom to do whatever it thinks is best, with no checklist.
The Surprise Finding:
The paper discovered that "freedom" doesn't always make the AI smarter.
- Some AIs (like the one named DeepSeek) were terrible when given total freedom but became excellent when forced to follow a checklist.
- Other AIs (like Claude) actually performed better when they had total freedom.
- The Lesson: You can't just say "AI is good" or "AI is bad." You have to know which AI works best with which style of management.
4. The "Real World" vs. "Simulation"
The paper ran a simulation for a year and also a "live" test for two weeks using a real broker (but with fake money).
- The Result: The "health check" worked perfectly in both. Even though the AI had never seen the live data before, the test could still tell you if it was calm, focused, or reliable.
- The "Gap": The paper found a consistent gap between what the AI said it would do and what it actually did. Even in the live test, the AI's actions often didn't match its reasoning. This suggests the AI is good at "talking the talk" but sometimes struggles to "walk the walk."
5. Why This Matters (The "Map" vs. The "Leaderboard")
The authors say: "Stop looking at the leaderboard."
Old tests tried to rank AIs from 1 to 10. This paper says that's useless because the market changes. An AI that wins today might lose tomorrow.
Instead, CLQT provides a map of limitations.
- It tells you: "This AI is great at staying calm but bad at following rules."
- It tells you: "This AI is smart but crashes if you don't give it enough time to think."
Summary
CLQT is a new, rigorous way to test AI investors. It stops them from cheating, forces them to explain their reasoning, and gives them a detailed report card on their personality and habits. It proves that making money in a simulation doesn't mean an AI is actually smart; it might just be lucky. The real value isn't in finding the "winner," but in understanding exactly where each AI is strong and where it is likely to fail.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.