← Latest papers
💰 quantitative finance

Tackling Decision Processes with Non-Cumulative Objectives using Reinforcement Learning

This paper introduces a general mapping that transforms Non-Cumulative Markov Decision Processes (NCMDPs) into standard MDPs, enabling the direct application of existing reinforcement learning techniques to optimize arbitrary reward functions and demonstrating improved performance and training efficiency across diverse tasks.

Original authors: Maximilian Nägele, Jan Olle, Thomas Fösel, Remmy Zen, Florian Marquardt

Published 2026-10-01
📖 6 min read🧠 Deep dive

Original authors: Maximilian Nägele, Jan Olle, Thomas Fösel, Remmy Zen, Florian Marquardt

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of artificial intelligence, there is a powerful framework used to teach machines how to make decisions. Imagine a robot learning to walk, a computer program mastering a video game, or a trading algorithm managing a stock portfolio. These systems operate by taking a series of actions, one after another, in response to their surroundings. With every move, the system receives a signal, often called a reward, telling it whether that action was good or bad. For decades, the standard rule for success in these scenarios has been simple: maximize the total sum of all rewards collected over time. If a robot gets a small point for every step forward, the goal is to get as many points as possible by the end of the journey. This approach, known as a Markov decision process, has been incredibly successful, guiding everything from industrial robots to self-driving cars.

However, real life is often more complicated than a simple tally sheet. Sometimes, the most important outcome isn't the total amount of good things that happened, but rather the worst moment that occurred, or the consistency of performance over time. Consider a spacecraft landing on a planet. The goal isn't just to land safely; it is to ensure the craft never exceeds a dangerous speed during the entire descent, regardless of how smooth the rest of the flight was. In finance, an investor might care less about the total profit made over a year and more about how much that profit fluctuated, seeking a steady return rather than a risky gamble. These scenarios involve what researchers call non-cumulative objectives, where the final score depends on a specific function of the entire history of rewards, such as the maximum value reached or the ratio of average gain to volatility. Until now, teaching artificial intelligence to optimize these complex, history-dependent goals has been difficult, often requiring custom-built algorithms that are hard to apply to new problems.

A team of researchers from the Max Planck Institute for the Science of Light and Friedrich-Alexander-Universität Erlangen-Nürnberg has developed a general solution to this problem. They discovered a way to translate these complex, non-cumulative challenges into the standard format that existing, powerful artificial intelligence tools already know how to solve. Instead of inventing a new type of learning algorithm from scratch, they created a bridge. They showed that by slightly changing how the machine perceives its current situation and how it calculates its immediate feedback, any complex goal can be converted into a standard "sum of rewards" problem. This allows researchers to take the most advanced, off-the-shelf learning software available today and apply it directly to problems that were previously out of reach, without needing to modify the software itself.

The core of their method involves giving the artificial agent a bit more memory. In a standard setup, an agent only needs to know its current state to make a decision. But when the goal depends on the entire history of rewards—like remembering the highest speed reached so far—the agent needs to carry that information with it. The researchers proposed a system where the agent's "state" is expanded to include a running summary of the past, such as the highest or lowest reward seen up to that moment. Simultaneously, they adjusted the immediate reward the agent receives at each step. Instead of getting a reward that simply reflects the current action, the agent receives a calculated value that, when added up over the whole journey, perfectly reconstructs the complex goal. For example, if the goal is to minimize the maximum speed, the agent is rewarded in a way that penalizes it only when it sets a new speed record, effectively turning the "minimum of the maximums" problem into a standard sum.

This approach was tested across a wide variety of difficult tasks, proving its versatility. In a simulation of a lunar lander, the researchers trained an agent to land a spacecraft while strictly limiting its maximum speed. They compared their method against a standard approach that tried to approximate the goal by adding a penalty at the very end of the flight. The new method, which treated the speed limit as a continuous part of the learning process, found a much better balance between landing safely and moving efficiently. In the realm of finance, they applied the technique to portfolio optimization, where the goal is to maximize the Sharpe ratio—a measure of risk-adjusted return that divides average profit by the volatility of those profits. Previous methods had to rely on rough approximations of this ratio. By using the new mapping, the agents could learn to maximize the exact ratio directly, resulting in significantly better investment strategies during training.

The researchers also explored discrete optimization problems, such as finding the most efficient arrangement of quantum logic gates or simplifying complex diagrams used in quantum computing. In these tasks, the goal is often to find the single best state reached during a long search, rather than the sum of all improvements made along the way. Here, the new method allowed the agents to explore more boldly. Because the agent was not penalized for temporary setbacks that were necessary to reach a better solution later, it learned faster and found higher-quality solutions than agents trained with standard cumulative rewards. In one experiment involving quantum error correction, the new method improved performance by a significant margin, finding better solutions in less time.

The strength of this work lies in its simplicity and generality. The researchers did not create a new learning algorithm; they created a translation layer. This means that any expert in a specific field, from robotics to finance, can take their existing problem, apply this mapping, and immediately use the most powerful reinforcement learning tools available. The method works in both predictable environments and those full of random noise, and it handles both simple and complex goals. While the researchers noted that the expanded memory required for the agent can make the problem slightly larger, modern deep learning techniques are well-equipped to handle this. The result is a unified framework that removes the barrier between complex, real-world objectives and the sophisticated tools of artificial intelligence, opening the door for machines to learn strategies that were previously too difficult to define.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →