Do VLMs Truly "Read" Candlesticks? A Multi-Scale Benchmark for Visual Stock Price Forecasting
This paper introduces a multi-scale benchmark dataset and evaluation framework to assess Vision-Language Models' ability to interpret candlestick charts for stock price forecasting, revealing that while these models perform well in persistent trends, they struggle with complex market scenarios and exhibit significant limitations in precise temporal reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a super-smart robot how to predict the stock market. You show it a picture of a stock chart (a "candlestick chart," which looks like a series of colorful bars showing price movements) and ask, "Will this stock go up or down in the next month?"
This paper asks a simple but tricky question: Is the robot actually "reading" the chart, or is it just guessing based on the text you wrote next to it?
Here is the story of their experiment, explained with some everyday analogies.
1. The Problem: The "Cheat Sheet" Confusion
In the past, researchers gave robots a mix of everything: the chart image, a news article, a spreadsheet of numbers, and a summary of the economy.
- The Analogy: Imagine taking a math test where you are allowed to use a calculator, a textbook, and a friend who whispers the answers in your ear. If you get an A, did you actually learn math, or did you just listen to your friend?
- The Issue: Because the robots had so much text to read, no one knew if they were actually understanding the visual patterns of the stock chart or just reading the text descriptions.
2. The Solution: The "Blindfold" Test
To fix this, the authors built a special test where the robots were blindfolded to text.
- The Setup: They gave the robots only the picture of the stock chart. No news, no numbers, no descriptions. Just the visual bars.
- The Multi-Scale Twist: They didn't just show one picture. They showed two:
- A Daily Chart (like looking at a street map to see traffic right now).
- A Weekly Chart (like looking at a highway map to see the long-term flow of traffic).
- The Goal: A good trader looks at both to make a decision. The researchers wanted to see if the robots could do the same thing: combine the "short-term" view with the "long-term" view to predict the future.
3. The Experiment: The Robot Race
They tested several famous AI models (like GPT-4, Claude, and Gemini) against a classic computer program called XGBoost (think of XGBoost as a very disciplined, old-school accountant who only looks at raw numbers).
They asked the robots to predict the stock's return over the next 30 days.
4. The Results: What Did They Find?
🚩 The "Trend Follower" Bias
The robots were great at spotting obvious trends.
- The Analogy: If a stock is going up like a rocket, the robot says, "It's going up!" If it's crashing like a stone, the robot says, "It's going down!"
- The Catch: They were terrible at the messy middle ground. When the market was sideways or confusing, the robots often got it wrong. They are like a weather app that is perfect at predicting "Sunny" and "Stormy" but terrible at predicting "Cloudy."
⏳ The "Short-Term Memory" Problem
The researchers asked the robots to predict 30 days into the future.
- The Surprise: The robots were actually better at predicting what would happen in 5 days.
- The Analogy: It's like asking a student to write a thesis on the next 30 years of history. Instead, they just wrote a really good summary of what happened yesterday. The robots are very good at spotting immediate patterns but struggle to plan for the distant future.
🎯 The "Optimist vs. Pessimist" Personalities
Every robot had a different personality:
- Claude was very cautious. It rarely predicted a rise unless it was 100% sure (High Precision), but it missed a lot of opportunities (Low Recall). It's like a conservative investor who only buys when it's safe.
- GPT-5 Mini was aggressive. It predicted rises often, catching more opportunities but making more mistakes. It's like a gambler who bets on everything.
- The Conclusion: There is no "perfect" robot. You have to pick the one that matches your risk tolerance.
5. The Big Takeaway: Why Pictures Matter
The most interesting finding was that the robots (Vision-Language Models) actually did better than the old-school number-crunching program (XGBoost), even though they were only looking at pictures.
- The Analogy: Imagine trying to learn a dance.
- The Old Way (XGBoost): You are given a list of numbers describing every step (Step 1: move foot 2 inches left, Step 2: move foot 3 inches right). You have to mathematically figure out the rhythm.
- The New Way (VLMs): You are shown a video of the dance. You can "see" the flow, the shape, and the rhythm instantly.
- The Result: The robots could "see" the shape of the market trends in the pictures much better than the computer could calculate them from numbers. The visual format acts like a shortcut for the brain.
Summary
This paper is a reality check for the stock market AI hype.
- Yes, AI can "read" charts, but it's not a crystal ball.
- It's great at spotting obvious trends (up or down) but bad at predicting the future when things are uncertain.
- It has a short attention span, focusing more on the immediate future than the long term.
- Visuals are powerful: Showing a robot a picture of a chart helps it understand the market better than just feeding it a spreadsheet of numbers.
The Bottom Line: These robots are excellent tools for a human trader to get a "second opinion" on short-term trends, but they shouldn't be trusted to make long-term investment decisions on their own just yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.