← Latest papers
💰 quantitative finance

UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos

This paper introduces UniFinEval, the first unified multimodal benchmark covering text, images, and videos across five core financial scenarios, which reveals that while current models like Gemini-3-pro-preview lead in performance, they still significantly lag behind financial experts in handling high-density information and complex reasoning.

Original authors: Zhi Yang, Lingfeng Zeng, Fangqi Lou, Qi Qi, Wei Zhang, Zhenyu Wu, Zhenxiong Yu, Jun Han, Zhiheng Jin, Lejie Zhang, Xiaoming Huang, Xiaolong Liang, Zheng Wei, Junbo Zou, Dongpo Cheng, Zhaowei Liu, Xin
Published 2026-02-02
📖 5 min read🧠 Deep dive

Original authors: Zhi Yang, Lingfeng Zeng, Fangqi Lou, Qi Qi, Wei Zhang, Zhenyu Wu, Zhenxiong Yu, Jun Han, Zhiheng Jin, Lejie Zhang, Xiaoming Huang, Xiaolong Liang, Zheng Wei, Junbo Zou, Dongpo Cheng, Zhaowei Liu, Xin Guo, Rongjunchen Zhang, Liwen Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help run a massive, high-stakes investment firm. This assistant needs to be a "super-analyst" who can read thick financial reports, interpret complex charts, watch hours of market analysis videos, and then make smart decisions about where to put money.

For a long time, we've been testing these AI assistants with simple quizzes—like asking them to read a single paragraph or identify a simple graph. But in the real world, finance is messy. It's like trying to navigate a stormy ocean while reading a map, listening to a radio broadcast, and watching a live video feed all at once. The old tests were too easy and didn't reflect the chaos of the real job.

Enter "UniFinEval": The Ultimate Financial Bootcamp

The authors of this paper created a new, super-challenging test called UniFinEval. Think of it as a "survival course" for AI models, designed specifically to see if they can handle the high-pressure, information-overload environment of real finance.

Here is how they built this bootcamp:

1. The Five Levels of the Bootcamp

Instead of just asking one type of question, the test is divided into five realistic scenarios, getting harder as you go:

  • Level 1: The Detective (Financial Statement Auditing): Imagine a detective looking at a messy crime scene. The AI has to find tiny inconsistencies in a financial report that is packed with text, charts, and confusing layouts. It's like finding a single typo in a 100-page document while someone is shouting distractions in your ear.
  • Level 2: The Investigator (Company Fundamental Reasoning): Now the AI has to figure out if a company is actually healthy. It needs to pull numbers from different reports, do complex math, and connect the dots between a text description and a graph. It's like solving a puzzle where the pieces are in different languages.
  • Level 3: The Trend Spotter (Industry Trend Insights): The AI zooms out. It's no longer looking at one company but an entire industry. It has to compare dozens of companies over time to spot big patterns. It's like watching a football game and predicting the winner based on the performance of every single player on the field, not just the star quarterback.
  • Level 4: The Radar (Financial Risk Sensing): This is the scary part. The AI has to watch a video of a market analyst talking while reading a report and looking at a chart. It needs to spot hidden dangers (risks) that are buried in the noise. It's like trying to hear a whisper in a hurricane.
  • Level 5: The General (Asset Allocation Analysis): The final boss level. The AI has to take everything it learned from the previous four levels and make a final decision on how to invest money, balancing all the risks and rules. It's the moment the General has to decide where to send the troops.

2. The Rules of the Game

  • No Cheating: The test questions weren't written by other AIs (which can make mistakes or lie). They were written by real human experts—people with advanced degrees and years of experience in finance.
  • The Mix: The test uses text, images, and videos together. You can't just read the answer; you have to "see" the chart and "watch" the video to understand the context.
  • The Size: They created nearly 4,000 of these tough questions in both English and Chinese.

3. The Results: Who Passed?

The researchers put 10 of the smartest AI models (the "students") through this bootcamp. Here is what happened:

  • The Top Performer: One model, Gemini-3-pro-preview, came out on top. It got about 74% of the answers right. That sounds good, but remember, this is a very hard test.
  • The Gap: Even the best AI still made a lot of mistakes compared to a human expert, who got about 90-95% right. The AI is good at reading the text, but it struggles when it has to connect the dots between a video, a chart, and a report.
  • The Struggles:
    • Math Trouble: Many models got the simple math wrong.
    • Hallucinations: Some models made up facts that weren't there (like a student guessing the answer because they were scared to say "I don't know").
    • The "Video" Problem: When videos were involved, most models got lost. They couldn't keep track of the story over time.
    • The Final Boss: In the hardest level (Asset Allocation), the AI's performance dropped significantly. They could gather the info, but they couldn't make the final, complex decision without getting confused.

4. The Big Takeaway

The paper concludes that while AI is getting very good at reading and understanding simple financial data, it is not yet ready to replace human financial experts in complex, real-world situations.

Think of it this way: The AI is like a brilliant student who can memorize a whole library of textbooks. But when you put them in a real war room with a live video feed, a shaking hand, and a ticking clock, they start to panic and make mistakes.

The authors built UniFinEval to show us exactly where the AI fails, so developers can fix those specific weaknesses before trusting these models with real money. They didn't say AI will soon be running the stock market; they said, "Here is a map of the cliffs, so we know where the AI is likely to fall off."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →