← Latest papers
📈 economics

DART: Domain Aware Relational Transformer for Stock Prediction

The paper proposes DART, a Domain-Aware Relational Transformer that combines a Multi-Scale Temporal Fusion module and a Domain-Knowledge-Guided Sparse Attention mechanism to achieve state-of-the-art stock prediction performance and significantly reduced computational costs on large-scale datasets.

Original authors: Ruoyu Huang, Xiaoyu Lin, Yibao Mao, Jiqiang Yan

Published 2026-08-06
📖 7 min read🧠 Deep dive

Original authors: Ruoyu Huang, Xiaoyu Lin, Yibao Mao, Jiqiang Yan

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to predict the weather, but instead of looking at clouds, you are looking at the stock market. This is the world of quantitative finance, where scientists and investors try to guess which stocks will go up or down tomorrow. For a long time, people thought the market was like a perfectly fair coin toss, where no one could ever really predict the future because everyone already knew everything. But then, smart people realized that humans aren't perfect robots; we get scared, excited, and follow the crowd, which makes the market a bit messy and actually predictable in some ways. Today, we use powerful computers and "deep learning" (a type of artificial intelligence that learns from patterns) to find these hidden signals in the noise. The big problem, though, is that there are thousands of stocks, and they all influence each other in complicated ways. Trying to figure out how every single stock talks to every other stock is like trying to listen to a conversation between every person in a stadium at once—it's too much information for a computer to handle without getting overwhelmed.

This is where a new study comes in with a clever solution called DART (Domain-Aware Relational Transformer). Think of the stock market as a giant, noisy party. In the past, if you wanted to know who was talking to whom, you might try to listen to every single conversation happening at the same time. That takes forever and leaves you with a headache. The researchers behind DART realized that people mostly talk to others in their own groups—like the kids at the snack table or the adults by the punch bowl. They don't usually shout across the room to someone in a completely different group. So, instead of listening to everyone, DART only listens to the people in the same "industry group" (like tech companies talking to other tech companies). This makes the computer's job much, much faster.

The paper suggests that by using this "group listening" strategy, combined with a special way of looking at time (checking for short-term trends, medium-term trends, and long-term trends all at once), the model can predict stock returns better than older methods. In tests using real data from the NASDAQ, NYSE, and S&P 500, the DART model showed it could make smarter investment choices and handle risk better than other top models. It also proved that by ignoring the "noise" of unrelated stocks, the computer could do its work using about 84% less energy than before. While the model isn't a magic crystal ball that guarantees profits, the authors suggest it is a much more practical and efficient tool for navigating the complex, crowded world of the stock market.

The Story of DART: A Smarter Way to Listen to the Market

The Problem: Too Much Noise, Too Many Connections

Imagine you are at a massive concert with thousands of people. If you want to know what song is playing, you could try to listen to every single person's voice at once. But that's impossible; it's just a wall of noise. This is exactly the problem computers face when trying to predict stocks. There are thousands of stocks, and every single one is connected to every other one. Old computer models tried to listen to everyone at the same time. This is called "full attention." It's very accurate in theory, but it's so slow and expensive that it's practically impossible to use for real-world trading with thousands of stocks. It's like trying to solve a puzzle by checking every single piece against every other piece; the puzzle gets too big, and your brain (or computer) crashes.

The Solution: The "Industry Neighborhood" Rule

The researchers proposed a simple but powerful idea: Stocks in the same industry are best friends; stocks in different industries are just acquaintances.

Think of the stock market as a giant city divided into neighborhoods. There's a "Tech Neighborhood," a "Food Neighborhood," and a "Car Neighborhood." People in the Tech Neighborhood talk to each other all the time because they share the same news and trends. But the people in the Tech Neighborhood rarely have a direct, urgent conversation with the people in the Food Neighborhood.

The DART model uses this rule. Instead of listening to the whole city, it only listens to the conversations happening within each neighborhood. This is called Domain-Knowledge-Guided Sparse Attention (DSA).

  • How it works: The computer looks at a stock, asks, "What neighborhood is this in?" and then only pays attention to the other stocks in that same neighborhood.
  • The Result: It ignores the useless noise from unrelated stocks. The paper shows this cuts the computer's workload by about 84% (specifically, from 0.23 GFLOPs down to 0.04 GFLOPs in their tests). It's like switching from trying to hear a whisper in a hurricane to just listening to a quiet chat at a coffee shop.

The Time Machine: Seeing the Past in Layers

Predicting the future isn't just about who is talking to whom; it's also about when things happen. Stocks move on different clocks. Some move fast (minute-by-minute), some move medium-fast (daily trends), and some move slow (monthly or yearly trends).

Old models often tried to look at time with just one pair of glasses, which meant they missed the big picture or the small details. DART uses a Multi-Scale Temporal Fusion (MSTF) module.

  • The Analogy: Imagine looking at a movie. You have a zoom lens that lets you see the tiny details of an actor's face (short-term), a wide lens that shows the whole scene (medium-term), and a drone shot that shows the whole city (long-term). DART uses all three lenses at the same time.
  • The Magic: It then uses a special "Progressive Temporal Fusion" network to stack these views together. It's like building a sandwich where every layer adds more flavor, helping the computer understand how a small change today might build up into a big trend tomorrow.

What They Found: Faster and Smarter

The researchers tested their new model, DART, against several other famous models (like LSTM, GCN, and even newer ones like MambaStock) using real data from the US stock market (NASDAQ, NYSE, and S&P 500) from 2020 to 2024.

Here is what the data suggests:

  1. Better Returns: DART didn't just guess better; it helped build a better "portfolio" (a collection of stocks). In tests, it achieved the highest Sharpe Ratio (a score that measures how much money you make compared to the risk you take) on the NASDAQ and S&P 500. For example, on the NASDAQ, DART scored a Sharpe Ratio of 2.432, while the next best was 1.827.
  2. Top Picks: When the model picked the top 10 stocks it thought would do best, it was right more often than the others (a metric called Precision@10). On the NYSE, it got 58.7% of its top picks right, beating the competition.
  3. The "No-Go" Zone: The study explicitly showed that if you remove the "neighborhood" rule and let the computer listen to everyone (Full Attention), the model gets confused and performs much worse. This proves that ignoring the irrelevant connections is actually good for the model.

The Bottom Line

The paper suggests that DART is a practical, efficient tool for the future of investing. It doesn't claim to be a magic wand that guarantees you'll get rich, but it does show that by using smart rules (like only listening to your own neighborhood) and looking at time in layers, computers can make much better sense of the chaotic stock market. It solves the "too much data" problem by being selective, making it possible to use powerful AI on thousands of stocks without breaking the bank or the computer.

As the authors note, there is still room to grow. Right now, the model only looks at "neighborhoods" based on static industry lists. In the future, they hope to make these neighborhoods dynamic, so the model can see when a tech company suddenly starts acting like a food company, or when news breaks that connects two totally different groups. But for now, DART stands as a strong, efficient step forward in teaching computers how to navigate the financial world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →