← Latest papers
🤖 AI

A Descriptive and Normative Theory of Human Beliefs in RLHF

This paper proposes and validates a new theoretical framework demonstrating that human beliefs about an AI agent's capabilities significantly influence preference generation in RLHF, showing that aligning these beliefs with actual agent capabilities—rather than assuming optimality—leads to more performant learned policies.

Original authors: Sylee Dandekar, Shripad Deshmukh, Frank Chiu, W. Bradley Knox, Scott Niekum

Published 2026-07-13
📖 6 min read🧠 Deep dive

Original authors: Sylee Dandekar, Shripad Deshmukh, Frank Chiu, W. Bradley Knox, Scott Niekum

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to drive a car. You show it two routes: a super-fast shortcut along a cliff edge, and a slower, safer road far from the edge. In the world of standard AI training, the human teacher is expected to pick the route that would be best if the robot were a perfect, super-genius driver who never made a mistake.

But here is the twist: real robots aren't perfect. They might wobble, they might misunderstand, and they might not be able to handle the tricky cliff edge.

This paper argues that when humans teach these robots, they don't just look at the map; they look at what they think the robot is capable of doing. If a human thinks, "This robot is a genius," they might pick the dangerous cliff shortcut. But if the robot is actually a bit clumsy, that shortcut could lead to a crash. If the human thinks, "This robot is a bit shaky," they might pick the safe, slow road, which is exactly what the robot needs to succeed.

The Big Idea: Beliefs Matter More Than You Think

The authors, a team of researchers from UMass Amherst and UT Austin, propose a new way to think about how humans give feedback to AI. They call this the "belief-based preference model."

Think of it like a coach talking to a rookie athlete.

  • The Old Way (The "Optimist" Coach): The coach assumes the rookie is already an Olympic champion. They say, "Go for the gold! Run the fastest, riskiest play!" If the rookie actually tries this, they trip and fall because they aren't ready.
  • The New Way (The "Realist" Coach): The coach looks at the rookie's actual skills. They say, "Okay, you're good, but you're still learning. Let's stick to the play that keeps you safe and gets you points."

The paper suggests that for AI to learn well, the human labelers (the coaches) need to have beliefs about the AI's skills that match reality. If the human thinks the AI is better than it actually is, the AI learns the wrong lessons and might fail spectacularly later.

What They Found (The Proof)

The team didn't just guess; they tested this in three different ways:

  1. The Math (The Theory): They did some heavy math to prove that if the human's belief about the robot's skills doesn't match the robot's actual skills, the robot will learn a "suboptimal" policy. In plain English: If the coach lies to the player about their skills, the player will play a bad game. They showed that the more the human's belief differs from reality, the worse the robot's performance gets.

  2. The Simulation (The Video Game Test): They built a digital world (a grid) where a robot had to navigate from a start point to a goal, avoiding "cliffs" that meant failure. They programmed the robot to be a little bit "noisy" (imperfect).

    • When the human "labeler" in the simulation believed the robot was perfect (zero noise), the robot learned to drive right next to the cliff and often fell off.
    • When the human believed the robot was imperfect (matching the robot's actual noise), the robot learned to take the safer, wider path and succeeded.
    • The Result: The best results happened when the human's belief matched the robot's actual capability. If the human thought the robot was perfect when it wasn't, the robot crashed.
  3. The Human Study (The Real People Test): This is the coolest part. They actually asked 146 real people to watch videos of a self-driving car and pick the better route.

    • The Setup: Before watching the videos, they "primed" the people.
      • Group A (Safe Priming): Watched a video of a car driving perfectly, obeying all laws, and being super safe.
      • Group B (Unsafe Priming): Watched a video of a car drifting, crashing, and driving recklessly.
      • Group C (Control): Watched nothing special.
    • The Result: The people who saw the "unsafe" video (Group B) started to believe the car was less capable. Because of this, they preferred the safer, slower routes in their choices. The people who saw the "safe" video (Group A) believed the car was a genius and preferred the riskier, faster routes.
    • The Stat: The difference between the "Safe" and "Unsafe" groups was statistically significant (a p-value of 0.015). This proves that what humans think the AI can do directly changes what they tell the AI to do.

What They Say Is NOT the Answer

The paper is very clear about what doesn't work.

  • It's not enough to just assume the AI is perfect. The standard way of training AI assumes humans pick the "best possible" path as if the AI were a god-like optimizer. The authors show that this is actually a bad idea if the AI is imperfect.
  • It's not just about the human being "irrational" or "biased" in a general sense. While other studies talk about human biases like "anchoring" or "groupthink," this paper focuses specifically on beliefs about the agent's capabilities. The paper argues that humans do not assume optimality when judging actions; instead, they rely on their specific beliefs about the agent's current capabilities. If these beliefs are incorrect (e.g., thinking the agent is perfect when it is not), the resulting preferences are misaligned, leading to poor outcomes.

The Takeaway for the Future

The authors suggest that if we want better AI, we need to fix the "mismatch" between what humans think the robot can do and what the robot can actually do.

They propose two simple ideas for the people building these systems:

  1. Tell the truth: Show the labelers the robot's actual limitations before they start grading. If the robot can't turn sharply, tell them that!
  2. Show, don't just tell: Before asking for feedback, show the labelers a video of the robot driving right now. If the robot is clumsy, show the clumsy driving. If it's good, show the good driving. This "primes" the human to have the right belief.

The paper concludes that this isn't a solved problem yet. They proved it works in simulations and with a small group of people, but they admit that doing this in a full, real-world AI training loop is a big next step. However, the message is clear: To teach a robot well, you have to know what the robot thinks it can do, and make sure the teacher knows the same thing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →