← Latest papers
⚡ electrical engineering

Exact Decomposition of Adversarial Dual-Objective Value Functions, with Applications to Optimal Drug Dosing

This paper establishes theoretical conditions under which exact decompositions of adversarial dual-objective value functions remain valid in Hamilton-Jacobi Reachability frameworks and demonstrates their application to solving optimal drug regimen design problems.

Original authors: Dylan Hirsch, William Sharpless, Sylvia Herbert

Published 2026-07-16
📖 8 min read🧠 Deep dive

Original authors: Dylan Hirsch, William Sharpless, Sylvia Herbert

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a spaceship navigating a chaotic asteroid field. You have a mission: get to a specific star (the goal), but you must never crash into an asteroid (the obstacle). Now, imagine there is a mischievous alien pilot trying to steer your ship into the rocks. In the world of robotics and safety engineering, scientists use a mathematical tool called "Hamilton-Jacobi Reachability" to figure out the perfect steering plan. Think of this tool as a super-smart GPS that doesn't just tell you the shortest path, but calculates the safest path that works no matter how the alien tries to mess you up. It turns the problem of "how do I survive?" into a giant, complex math puzzle called a "value function." This function acts like a weather map for your journey: if the number is positive, you can make it; if it's negative, you're doomed.

For a long time, this GPS was great at simple missions: "Get to the star" or "Stay away from rocks." But real life is messy. Sometimes you need to do two things at once, like "Get to the star, but also make sure you never get too close to the rocks, even after you arrive." Or, "Visit Star A and Star B, in any order you want." Recently, scientists found a clever trick to break these complex two-part missions into smaller, easier puzzles. However, there was a catch: this trick only worked when the alien pilot wasn't there. The moment you added a mischievous adversary, the math broke, and the old tricks stopped working. This left engineers stuck, unable to use their best tools for the most dangerous, real-world scenarios.

This paper steps in to fix that broken math. The authors, Dylan Hirsch, William Sharpless, and Sylvia Herbert, prove that those clever "decomposition" tricks actually do work, even when a sneaky adversary is trying to ruin the plan. They showed that you can still break down complex, two-part safety missions into simpler pieces, calculate the safety of each piece separately, and then stitch them back together to get the perfect, robust plan. They didn't just guess; they provided a rigorous mathematical proof that these shortcuts are exact and reliable in continuous time. To show off their new theory, they applied it to a life-or-death scenario: designing the perfect drug dosage for a patient. They demonstrated that their method could find a treatment plan that cures a disease without accidentally poisoning the patient's kidneys, even when the body's internal chemistry is unpredictable and "adversarial."

The Core Discovery: Taming the Chaos

The main finding of this work is that specific ways of breaking down complex safety problems—called "value function decompositions"—remain valid even when an adversary is present. In the world of control theory, an "adversary" is a mathematical representation of uncertainty or a malicious force trying to push the system into failure. The authors proved that for two specific types of complex missions, known as Reach-Always-Avoid (RAA) and Reach-Reach (RR), you can still use the "divide and conquer" strategy.

The RAA problem is like a mission where you must reach a target, but you must always avoid a danger zone, even after you've reached the target. The RR problem is like a scavenger hunt where you must visit two different locations, but you can visit them in whichever order you prefer.

The paper explicitly rules out the idea that these decompositions fail in the presence of an adversary. In fact, the authors provide a counter-example to show why a different, seemingly logical way of breaking down the problem (specifically for the "Reach-Reach" task) fails when an adversary is involved. They showed that if you try to simply choose the best order to visit targets based on a simple calculation, a clever adversary can force the system into a situation where that order fails, even if the mission is actually possible. This proves that you cannot just use the old "no-adversary" logic; you need the specific, new mathematical structures they developed.

The authors are extremely confident in these results. They didn't just simulate them; they provided formal mathematical proofs (Theorem 1 and Theorem 2) showing that these decompositions are exact. This means the math is not an approximation or a "good guess"; it is a precise equality. They established these results in a continuous-time setting, which is the standard for real-world physics and engineering, rather than a simplified "step-by-step" (discrete-time) world often used in computer games or basic reinforcement learning.

How It Works: The Magic of Splitting the Puzzle

To understand the magic, imagine you are trying to navigate a maze while a ghost tries to push you into walls.

The Reach-Always-Avoid (RAA) Mission:
Imagine you need to reach a treasure chest (Target) but must never touch the spikes (Obstacle). The old way of thinking said, "Just reach the chest while avoiding spikes." But the new RAA rule says, "Reach the chest, and then keep avoiding the spikes forever."
The paper shows you can solve this by doing two simpler things:

  1. First, calculate the "Avoid Value": How safe is it to stay away from the spikes, ignoring the treasure?
  2. Second, create a "New Treasure Map." This map says the treasure is only "real" if you are in a spot where you can reach it and stay safe from the spikes forever.
  3. Finally, solve the standard "Reach-Avoid" problem using this new map.
    The authors proved that the result of this three-step process is exactly the same as solving the giant, scary RAA problem all at once.

The Reach-Reach (RR) Mission:
Now imagine you have two treasure chests, Chest A and Chest B. You need to open both. You can go A then B, or B then A.
The paper shows you can solve this by:

  1. Calculating how easy it is to reach Chest A.
  2. Calculating how easy it is to reach Chest B.
  3. Creating a "Super Treasure" that is a combination of these two. This Super Treasure is found if you can reach Chest A and then Chest B, OR reach Chest B and then Chest A.
    The authors proved that solving for this "Super Treasure" gives you the exact answer for the complex RR problem, even if a ghost is trying to push you away from the chests.

Real-World Application: Saving Lives with Math

The authors didn't stop at theory; they showed how this math can save lives in optimal drug dosing.

Example 1: The Kidney Problem
In this scenario, a patient needs a drug to cure an illness (the "Reach" part), but the drug is toxic to the kidneys (the "Avoid" part).

  • The Problem: Traditional methods might give a huge dose to cure the patient quickly. This works for the cure, but the drug lingers in the blood and eventually floods the kidneys, causing toxicity. Even if you stop the drug the moment the cure is achieved, the drug already in the blood keeps flowing to the kidneys.
  • The Solution: Using the new RAA decomposition, the computer calculates a dosing schedule that reaches the cure threshold while ensuring the kidney concentration never crosses the toxic line, even after the treatment stops.
  • The Result: In their simulations, the traditional method led to kidney toxicity (the "dash-dot" and "dotted" lines in their graphs), while the new RAA method kept the patient safe (the "solid" line). The simulation used a model where the drug concentration in the blood (x1x_1) and kidneys (x2x_2) were tracked, with a toxic threshold of 1.0. The new method successfully kept the kidney concentration below 1.0 while the blood concentration reached the therapeutic goal.

Example 2: The Protein Balancing Act
In a second example, the goal was to increase the levels of two different proteins in a cell to fight a disease.

  • The Problem: If you try to boost both proteins at the same time, the cell's natural chemistry (which acts like an adversary) might cancel them out, and neither reaches the needed level.
  • The Solution: The RR decomposition allows the controller to time the production perfectly. It might boost Protein 1 first, wait for the cell to adjust, and then boost Protein 2.
  • The Result: The simulation showed that a "simultaneous" approach failed to reach the targets, but the RR approach successfully coordinated the timing to hit both therapeutic thresholds.

Why This Matters

This paper is a bridge between elegant math and messy reality. For years, engineers had to choose between using powerful, simple math tricks (that only worked in a perfect, no-adversary world) or using complex, slow, and often inaccurate methods for the real world. This work proves that you can have the best of both worlds: the simplicity of breaking a big problem into small ones, combined with the robustness needed to handle the worst-case scenarios.

The authors note that while they have cracked the code for these two specific types of missions, the door is now open to apply this logic to even more complex tasks, like those described by "signal temporal logic" (which can describe very intricate rules like "visit A, then avoid B, then visit C, but only if D happens"). They acknowledge that future work is needed to see which other "rules" of the game hold up when an adversary is playing, and to ensure that these mathematical switches between different control strategies work smoothly in real-world learning algorithms. But for now, they have firmly established that for reaching targets while avoiding danger, and for visiting multiple goals, the math holds up, even in the face of chaos.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →