A Reinforcement Learning Supervisor with Dynamic Performance-Metric Weighting for Cryptocurrency Portfolio Management
This paper proposes a reinforcement learning supervisor framework that dynamically weights multiple performance metrics to select the optimal trading policy from a heterogeneous set of agents, demonstrating superior risk-adjusted returns in nonstationary cryptocurrency markets compared to individual agents and fixed strategies.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Investing in financial markets is a constant game of adaptation. The rules of the game change as the economy shifts from a booming period of growth to a slow, sideways drift, or a sudden, sharp decline. In these environments, a single strategy that works perfectly during a boom often fails miserably when the market turns. This is the core challenge of portfolio management: deciding how to split money among different assets when the future is uncertain and the past is not a reliable map. For decades, investors have relied on mathematical formulas to balance risk and reward, but these static models often struggle when market conditions change rapidly. More recently, computer scientists have turned to a type of artificial intelligence called reinforcement learning. Instead of being taught specific rules, these computer programs learn by trial and error, interacting with a simulated market to discover which actions lead to the best long-term results. While these programs can learn to trade, they often get stuck in a single way of thinking, unable to pivot when the market regime changes.
A team of researchers at the Korea National University of Transportation has proposed a new way to solve this problem. Rather than trying to build one perfect trading robot, they created a system that acts like a supervisor, managing a team of five different trading robots, each trained with a slightly different learning method. The supervisor does not make the trades itself. Instead, it watches how each robot has performed recently and decides which one should be in charge at any given moment. The researchers tested this system using real, high-frequency data from the cryptocurrency market, a place known for its extreme volatility and rapid changes. They found that this supervisor system consistently outperformed any single robot, as well as traditional investment strategies, by dynamically switching to the agent best suited for the current market conditions.
The researchers started by building a team of five distinct agents. Each agent is a computer program designed to make investment decisions, but they were trained using different mathematical approaches. Although all agents were trained in the same portfolio environment with the same reward function, their distinct learning mechanisms caused each agent to develop a unique personality. One might be very aggressive, chasing high profits but taking big risks, while another might be more cautious, prioritizing stability over rapid gains. In a stable market, the aggressive agent might shine. In a turbulent market, the cautious one might be the only one that survives. The problem is that no single agent is the best at everything all the time. If an investor picks just one agent and sticks with it, they are likely to suffer when the market shifts in a way that agent dislikes.
To solve this, the researchers introduced a supervisor agent. This supervisor does not know the internal code of the five agents it is managing. Instead, it acts like a coach watching a game, looking only at the scoreboard. It tracks seven different performance metrics for each agent, such as how much profit they made, how much risk they took to get that profit, and the proportion of periods with positive returns. It looks at these numbers over different time periods, from the last few hours to the last few days. The supervisor's job is to figure out which of these metrics matters most right now. In a calm market, the supervisor might decide that steady, consistent profits are the most important sign of a good agent. In a chaotic market, it might decide that avoiding large losses is the only thing that counts.
The supervisor learns this weighting system through its own training process. It observes the market and the agents' recent histories, then learns to assign importance scores to the different performance metrics. It does not simply pick the agent with the highest recent profit. Instead, it calculates a score for each agent by combining their performance data with the importance weights the supervisor has learned for that specific moment. The agent with the highest total score is selected to manage the portfolio for the next trading period. This creates a system that is constantly re-evaluating its choices, shifting its trust from one agent to another as the market environment changes.
The researchers tested this system using thirty-minute price data from the Binance cryptocurrency exchange, covering a period from January 1, 2021, to December 31, 2025. They set up two different scenarios: one with four popular cryptocurrencies and another with five, adding a fifth asset to see if the system could handle a slightly larger group of investments. They compared their supervisor system against the five individual agents, a simple strategy of buying and holding all assets equally, and several traditional investment formulas. The results were clear. The supervisor system generated significantly higher final portfolio values than any of the individual agents or the traditional strategies. In the four-asset scenario, the supervisor grew the portfolio value to 2.349 times the starting amount, while the best individual agent only reached 1.652. In the five-asset scenario, the supervisor reached 3.044 times the starting value, compared to 2.554 for the best individual agent.
Crucially, this extra growth did not come at the cost of higher risk. The supervisor system also experienced smaller maximum losses, known as drawdowns, than the individual agents. This suggests that the system was not just taking bigger gambles to get higher returns; it was actually managing risk better by knowing when to switch to a safer agent. The researchers also compared their system to a simpler version that used only one fixed rule to switch agents, such as always picking the agent with the highest recent profit. The supervisor system vastly outperformed these single-rule strategies. This proved that the system's success came from its ability to weigh multiple factors simultaneously, rather than relying on a single, rigid criterion.
To understand exactly what made the system work, the researchers conducted a series of tests where they removed parts of the system to see what happened. They found that the system needed all seven performance metrics and all four time windows to work at its best. Removing the profit-related metrics hurt the system's ability to grow wealth, while removing the risk-related metrics hurt its ability to protect against losses. Similarly, the system needed to look at both short-term and long-term performance history. In the four-asset test, the long-term history was most important, while in the five-asset test, the short-term history mattered more. This variation showed that the system was not using a fixed rule but was genuinely adapting to the specific needs of the current portfolio and market conditions.
The study concludes that managing a portfolio is not just about finding the single best trading algorithm. It is about creating a structure that can choose the right tool for the job as the job changes. By using a supervisor to dynamically weigh different performance signals, the system can navigate the unpredictable nature of cryptocurrency markets more effectively than any single agent or static strategy. The researchers note that their current system uses a specific set of five agents and a specific set of market data. Future work could expand this by including agents with different risk profiles or testing the system on other types of assets. However, the core finding remains robust: a hierarchical approach that learns to select policies based on a flexible, multi-faceted view of performance offers a powerful way to handle the uncertainty of financial markets.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.