← Latest papers
💻 computer science

Risk-Aware Preference Learning for Stochastic Outcomes

This paper demonstrates that incorporating Cumulative Prospect Theory to model human risk sensitivity significantly improves reward function recovery and reduces regret in preference learning for stochastic robot navigation compared to traditional expected utility assumptions.

Original authors: Yi-Shiuan Tung, Yuni Wu, Wei Jiang, Alessandro Roncone, Bradley Hayes

Published 2026-07-20
📖 4 min read☕ Coffee break read

Original authors: Yi-Shiuan Tung, Yuni Wu, Wei Jiang, Alessandro Roncone, Bradley Hayes

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to behave like a polite human. You can't just write a long list of rules like "don't hit people" or "walk fast," because life is messy and full of surprises. Instead, scientists have found a clever trick: show the robot two different ways of moving and ask a human, "Which one looks better?" By doing this thousands of times, the robot can learn a hidden "scorecard" of what humans actually value. This is called preference-based reward learning.

However, there is a tricky assumption hiding in the background of most of these studies. It assumes that humans are like perfect calculators who simply add up the good and bad parts of a situation, weighing them by how likely they are to happen. This is called Expected Utility. But real humans aren't calculators; we are emotional. We worry way too much about rare disasters (like a car crash) and we hate losing things more than we love gaining them. This paper asks a simple question: What happens if we teach a robot using a model that assumes humans are perfect calculators, when the humans are actually risk-averse, emotional decision-makers?

The researchers at the University of Colorado Boulder decided to test this idea in a simulated world where a robot has to navigate through a crowd of pedestrians. In this world, every move the robot makes isn't a single, guaranteed path. Instead, it's like rolling a dice: the robot might glide smoothly, or it might bump into someone, or it might get stuck. The outcome is a cloud of possibilities, not a straight line.

The team set up a digital experiment with two types of "teachers" (the humans whose preferences the robot tries to learn). One teacher was a "perfect calculator" who only cared about the average outcome (Expected Utility). The other teachers were "risk-sensitive" humans who used a complex psychological model called Cumulative Prospect Theory (CPT). This model accounts for how real people overestimate the chance of rare, scary events and feel the pain of a collision much more sharply than the joy of a smooth path.

The researchers then trained two types of "students" (the learning algorithms). One student assumed the teacher was a perfect calculator (the EU learner), while the other student assumed the teacher was a risk-sensitive human (the CPT learner). They fed both students thousands of "which path is better?" questions generated by the teachers and watched how well they learned the true scorecard.

Here is what they found in their simulations. When the teacher was a perfect calculator, the calculator-student learned perfectly, while the risk-sensitive student got a little confused by the extra complexity. But when the teacher was a risk-sensitive human, the calculator-student performed significantly worse. It tried to explain the human's fear of rare crashes by distorting the recovered scorecard, making the robot think collisions were just "a little bit bad" instead of "catastrophic," resulting in a biased reward. The robot learned a less accurate lesson.

In contrast, the risk-sensitive student (the CPT learner) performed much better. It realized that the human wasn't just valuing the outcome differently; they were weighing the uncertainty differently. By separating the "value" of the outcome from the "fear" of the risk, the CPT learner recovered a scorecard with substantially lower error than the calculator-student.

The paper suggests that if we want robots to be safe and aligned with human values in the real world—where rare accidents are scary and people are naturally anxious about them—we can't just assume humans are logical math machines. We need to build robots that understand human fear and risk sensitivity, or else they might learn to take dangerous shortcuts because they think we don't mind the risk. While this result comes from computer simulations and not yet from real people testing real robots, the evidence points to a clear need: to teach robots how to be safe, we must first teach them how humans actually feel about the unknown.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →