← Latest papers
💬 NLP

Herculean: An Agentic Benchmark for Financial Intelligence

The paper introduces Herculean, the first agentic benchmark for financial intelligence that evaluates AI agents across four complex workflows (Trading, Hedging, Market Insights, and Auditing) using standardized MCP-based environments, revealing that while agents perform well on trading and insights, they significantly struggle with the long-horizon coordination and structured verification required for hedging and auditing.

Original authors: Xueqing Peng, Zhuohan Xie, Yupeng Cao, Haohang Li, Lingfei Qian, Yan Wang, Vincent Jim Zhang, Huan He, Xuguang Ai, Linhai Ma, Ruoyu Xiang, Yueru He, Yi Han, Shuyao Wang, Yuqing Guo, Mingyang Jiang, Yi
Published 2026-05-15
📖 5 min read🧠 Deep dive

Original authors: Xueqing Peng, Zhuohan Xie, Yupeng Cao, Haohang Li, Lingfei Qian, Yan Wang, Vincent Jim Zhang, Huan He, Xuguang Ai, Linhai Ma, Ruoyu Xiang, Yueru He, Yi Han, Shuyao Wang, Yuqing Guo, Mingyang Jiang, Yilun Zhao, Youzhong Dong, Xiaoyu Wang, Yankai Chen, Ye Yuan, Qiyuan Zhang, Fuyuan Lyu, Haolun Wu, Yonghan Yang, Zichen Zhao, Yuyang Dai, Fan Zhang, Rania Elbadry, Ayesha Gull, Muhammad Usman Safder, Nuo Chen, Fengbin Zhu, Tianshi Cai, Zimu Wang, Polydoros Giannouris, Yuechen Jiang, Zhiwei Liu, Mohsinul Kabir, Yuyan Wang, Yixiang Zheng, Yangyang Yu, Weijin Liu, Wenbo Cao, Anke Xu, Peng Lu, Jerry Huang, Fengran Mo, Mingquan Lin, Prayag Tiwari, Yijia Zhao, Victor Gutierrez Basulto, Xiao-Yang Liu, Kaleb E Smith, Jiahuan Pei, Arman Cohan, Jimin Huang, Yuehua Tang, Alejandro Lopez-Lira, Xi Chen, Xue Liu, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new financial assistant. In the past, you might have tested them by asking, "What was Apple's stock price yesterday?" or "Summarize this earnings report." If they got the answer right, you thought they were smart.

But the authors of this paper argue that getting a single answer right isn't the same as being able to do a real job. They created a new test called HERCULEAN (a name playing on the idea of a "Herculean" or massive task) to see if AI agents can actually work like professional financial experts, not just answer trivia questions.

Here is a breakdown of their findings using simple analogies:

The Four "Jobs" They Gave the AI

Instead of a simple quiz, the researchers gave the AI four distinct, complex jobs that real humans do every day. Think of these as four different shifts in a financial factory:

  1. The Trader (Trading):

    • The Job: Every day for three months, the AI has to decide whether to buy, sell, or hold a specific stock.
    • The Catch: It's like playing a game of chess where the board changes every day, and you can't see the future. You have to make a move, live with the consequences, and then make the next move based on only what you know right now.
    • The Result: The AI was okay at this. It could make decisions, though not always profitable ones.
  2. The Risk Manager (Hedging):

    • The Job: This is like a tightrope walker. The AI has to pick two stocks that move in opposite directions (or at least differently) and bet on the gap between them, not on the market going up or down.
    • The Catch: It requires remembering the relationship between two different things over a long time. If you forget the details of the first stock while looking at the second, you fall off the tightrope.
    • The Result: The AI struggled here. It had trouble keeping track of the long-term relationship between the two assets.
  3. The Analyst (Market Insights):

    • The Job: The AI has to read news, prices, and reports, then write a professional investment report every week recommending whether to buy or sell.
    • The Catch: It's not just about finding facts; it's about weaving them into a coherent story that makes sense to a human investor.
    • The Result: The AI was surprisingly good at this. It wrote fluent, well-structured reports that looked very professional.
  4. The Auditor (Auditing):

    • The Job: The AI has to act like a strict accountant. It looks at a company's official financial documents and checks if the numbers add up correctly according to complex government rules.
    • The Catch: There is no room for "maybe" or "it looks right." It's a math puzzle where if you miss one tiny decimal point or misinterpret a rule, the whole thing is wrong.
    • The Result: This was the hardest job. The AI failed frequently. It often couldn't follow the strict rules or do the math correctly, even though it was good at writing stories.

The Big Discovery: "Fluency" vs. "Reliability"

The paper found a strange gap in how the AI thinks.

  • The "Storyteller" Problem: The AI is great at fluent generation. If you ask it to write a report or explain a trend, it sounds confident and smart. It's like a student who can write a beautiful essay but fails the math test.
  • The "Executor" Problem: The AI is bad at structured verification. When the job requires strict logic, checking facts against rules, or remembering a state over a long time (like the Auditor or Risk Manager jobs), it falls apart.

The "Toolbox" Analogy

The researchers also tested different "frameworks" (the software shells the AI runs inside). They found that the AI's performance depended heavily on how it was allowed to use its tools.

  • The "Lightweight" Agent: Imagine an AI that tries to do everything in its head without writing things down. It gets confused easily, especially in long tasks.
  • The "Structured" Agent: Imagine an AI that is forced to use a checklist, write down every step, and verify its work before moving on. This version performed much better.

The Lesson: It's not just about having a "smarter" brain (a better AI model); it's about having a better system to manage that brain. A smart brain without a good system to keep it on track will still make mistakes in high-stakes jobs.

The Conclusion

The paper concludes that while AI is getting better at talking about finance, it is not yet ready to do the job of a professional financial worker. It can write the report (Market Insights), but it struggles to manage the risk (Hedging) or verify the math (Auditing).

The researchers built HERCULEAN to show us exactly where these "muscles" are weak, so developers can build better systems that don't just sound smart, but actually work reliably.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →