An Inter-variable Relationship-Aware Mixture-of-Experts Model for Stock Closing-Price Prediction
The paper proposes IR-MoE, an Inter-variable Relationship-Aware Mixture-of-Experts Transformer that explicitly models dynamic dependencies among trading variables and employs sparse routing with a stock-wise shuffled training strategy to achieve superior forecasting accuracy and zero-shot generalization across diverse global stock markets.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Predicting the future price of a stock is one of the most persistent challenges in finance. The market is a chaotic place, influenced by everything from global news and government policies to the collective mood of millions of investors. Because of this, stock prices do not move in simple, straight lines; they are noisy, erratic, and constantly shifting. For decades, experts have tried to find patterns in this chaos, often looking at a standard set of five daily numbers: the price when the market opens, the highest point it reached, the lowest point, the price when it closes, and the total volume of shares traded. The difficulty lies not just in tracking these numbers over time, but in understanding how they talk to each other. A rise in price accompanied by heavy trading volume tells a different story than a rise with little volume, yet many computer models treat these five numbers as separate, unrelated streams of data.
The core question researchers face is whether a single computer program can learn the complex, shifting rules of the stock market well enough to apply them to new, unseen stocks without needing to be retrained from scratch. Traditional methods often struggle here, either failing to capture the subtle relationships between price and volume or becoming so specialized to one specific company that they cannot adapt to another. This limitation makes it difficult to build a universal tool for investment analysis that works across different countries and market conditions.
In a recent study, a team of researchers from the Beijing Institute of Technology proposed a new approach to solve this problem. They developed a system called IR-MoE, which stands for Inter-variable Relationship-Aware Mixture-of-Experts. The name describes exactly what the system does: it is designed to be aware of how the different trading variables relate to one another, and it uses a "mixture of experts" to make predictions. Instead of forcing the computer to learn a single, rigid rule for all stocks, this system allows different parts of the model to specialize in different market conditions. It is built on the idea that while every stock is unique, the way price and volume interact often follows recognizable patterns that can be learned and transferred.
The researchers began by gathering a massive amount of historical data from four distinct financial markets: the Chinese A-share market, the United States, Hong Kong, and Japan. They selected hundreds of representative stocks from these regions, covering a wide range of industries and market behaviors. Rather than training a separate model for each of these hundreds of stocks, they fed all the data into a single, unified system. This "union training" strategy allowed the model to learn shared market dynamics, such as how a surge in trading volume often signals a shift in price direction, regardless of whether the stock is based in Shanghai or New York.
The heart of their system is a two-part architecture. The first part focuses on understanding the relationships between the five daily numbers. The researchers treated each number—open, high, low, close, and volume—as a distinct piece of information, or token, and used a mechanism inspired by modern language models to let them interact. This allowed the system to explicitly learn, for example, that the closing price is heavily influenced by the trading volume of that day. By visualizing these interactions, the researchers found that the model consistently paid close attention to the volume when trying to understand price movements. This attention was not random; it was stronger during periods of high volatility, suggesting the model had learned that volume becomes a more critical clue when the market is turbulent.
The second part of the system is the "mixture of experts." Imagine a team of specialists, each an expert in a specific type of market behavior, such as a steady upward trend, a sudden crash, or a quiet, sideways movement. When the system receives data for a new stock, it does not ask every specialist to weigh in. Instead, a routing mechanism looks at the current market conditions and selects only a few of the most relevant experts to make the prediction. This allows the system to adapt instantly to different situations. If the market is behaving erratically, it might call upon experts trained on high-volatility patterns; if the market is calm, it might rely on those trained on steady trends. This flexibility is what allows the model to handle the diverse and changing nature of global stock markets.
The results of the study were tested across all four markets, including a rigorous test where the model was asked to predict prices for stocks it had never seen before, without any additional training. In these "zero-shot" tests, the new system outperformed several existing methods, including complex models that break data into smaller pieces and other advanced artificial intelligence systems designed for time-series forecasting. On the Chinese A-share market, the model achieved a prediction error rate of about 1.10 percent, while on the U.S. market, it reached an error rate of roughly 1.32 percent. These figures were consistently better than the best-performing alternatives in the study.
Perhaps most significantly, the model proved it could transfer its knowledge across borders. When the system trained on data from Chinese stocks was applied directly to unseen stocks in the United States, Hong Kong, and Japan, it maintained high accuracy. This suggests that the fundamental relationships between price and volume are universal enough to be learned once and applied anywhere. The researchers also found that the system worked best when it was trained on a diverse but manageable number of stocks. Adding more data helped at first, but beyond a certain point, the benefits diminished, indicating that the key was covering a wide variety of market behaviors rather than simply feeding the system more numbers.
To ensure the system was working correctly, the researchers also examined how the "experts" were being used. They found that without a specific balancing mechanism, the system tended to rely too heavily on a few experts, leaving others idle. By adding a rule that forced the system to distribute the work more evenly, they ensured that all the specialized parts of the model were active and learning. This attention to detail in the training process was crucial for the system's stability and performance.
The study concludes that by explicitly modeling how trading variables interact and by using a flexible team of specialized predictors, it is possible to build a more robust and adaptable tool for stock forecasting. The researchers suggest that their approach offers a promising path forward for financial analysis, one that moves away from rigid, single-purpose models toward systems that can understand the complex, shifting language of the market. While the study does not claim to have solved the mystery of the stock market, it demonstrates that a deeper understanding of the relationships between daily trading data can lead to more accurate and generalizable predictions. The work opens the door for future research that could incorporate even more types of data, such as news reports or intraday trading details, to further refine our ability to navigate the financial world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.