← Latest papers
🤖 AI

A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry

This paper introduces a contextual-bandit oversight game model featuring two-sided informational asymmetry to characterize the "avoidable harm" gap between optimal team performance and myopic human behavior, analyzing how this inefficiency can be dynamically resolved through repeated interaction and signaling.

Original authors: Yunjin Tong

Published 2026-07-02
📖 5 min read🧠 Deep dive

Original authors: Yunjin Tong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a high-stakes game of "Guess What I Know" played between a human boss and a smart robot assistant. This paper studies a specific problem: What happens when both of them are hiding something important from each other?

Usually, we think of AI problems as the robot not knowing what the human wants (e.g., "Does the boss want speed or safety?"). But this paper flips the script. Here, the human knows their own preferences (like "I hate it when things break"), but the robot knows something the human can't see (like "The shelf I'm about to grab is actually cracked").

The paper asks: When should the robot ask for help, and when should the human step in to stop the robot?

The Setup: The Warehouse Robot and the Floor Manager

To make this concrete, the authors use an analogy of a warehouse:

  • The Human (The Manager): She stands on the floor. She knows she gets fired if a shelf collapses (safety) but gets a bonus if items move fast (speed). She doesn't know the technical details of the robot's sensors.
  • The Robot (The AI): It has high-tech cameras and force sensors. It can "feel" that a shelf bracket is weak and will collapse if it grabs an item quickly. The human cannot see this.
  • The Interaction: The robot proposes an action ("I will grab this item fast"). The human has three choices:
    1. Trust: Let the robot do it.
    2. Ask: "Wait, are you sure?" (This costs time/money).
    3. Oversee: "Stop! I'm taking over." (This also costs time/money).

The Core Problem: The "Slab" of Avoidable Harm

The paper finds a dangerous gap in how humans and robots usually interact. They call this gap a "Slab of Avoidable Harm."

Here is how the failure happens in everyday terms:

  1. The Robot's Secret: The robot sees the shelf is cracked (Harmful). It knows that if it grabs the item, the shelf will fall.
  2. The Human's Guess: The human looks at the situation and thinks, "Statistically, shelves usually hold up. The robot is probably fine." She trusts her gut feeling (her "prior").
  3. The Misunderstanding: The robot could ask for help, but it knows the human is likely to say "No, go ahead" because she trusts her gut. So, the robot stays silent to save time.
  4. The Disaster: The robot grabs the item, the shelf falls, and the warehouse is a mess.

The Tragedy: The robot knew it was dangerous, and the human would have stopped it if she had known the truth. But because the robot didn't ask (thinking it wouldn't help) and the human didn't override (trusting her guess), the disaster happened.

The Two Solutions Compared

The paper compares two ways this game could be played:

1. The "Myopic" (Short-Sighted) Way:
This is what happens in the real world right now. The human treats the robot's questions as just noise. If the robot asks, she thinks, "It's probably fine, but I'll check my stats." If her stats say "It's probably fine," she ignores the robot.

  • Result: The "Slab" of harm exists. The robot stays silent, the human stays confident, and things break.

2. The "Team Optimal" (Perfect Coordination) Way:
Imagine the human and robot agree on a secret code beforehand: "If the robot asks, it means 'I see a disaster coming.' If you hear an ask, you MUST stop immediately."

  • Result: The robot asks whenever it sees danger. The human stops immediately. The disaster is avoided.
  • The Catch: This only works if the human believes the robot is telling the truth. If the robot asks just to be annoying, the human stops listening.

The "Price of Non-Credible Communication"

The paper argues that the "Slab of Harm" is the price we pay for not having a credible signal.
If the robot's "Ask" button is a credible signal (meaning "I know something bad is happening"), the human ignores her own guess and listens. If the "Ask" button isn't credible, the human sticks to her guess, and the robot stays silent.

How to Fix It Over Time

The paper also looks at what happens if they play this game many times (like a robot working in a warehouse for months). Even if the human is short-sighted at first, two things can fix the problem:

  1. Passive Learning (The "Burned Finger" Effect):
    If the robot keeps trying to grab items and shelves keep falling, the human eventually learns: "Oh, when this robot tries to grab fast, things break." She updates her belief. Eventually, she stops trusting the robot blindly and starts listening to it. It takes time and some accidents, but she learns.

  2. Active Signaling (The "Crying Wolf" Strategy):
    The robot can force the issue. Even if the human usually ignores it, the robot can start asking every single time it sees a danger.

    • The Cost: The robot has to pay a "cost" (time/money) to ask.
    • The Payoff: If the human sees the robot asking constantly, she realizes, "Wait, this robot is only asking when things are bad." She updates her belief instantly. The robot pays a small cost now to save a huge disaster later.

Summary

The paper proves that when an AI knows something dangerous that a human can't see, and the human doesn't trust the AI's warnings, harm is inevitable.

The solution isn't just "make the AI smarter." It's about designing the system so that asking for help is a credible signal. If the human knows that "Asking = Danger," she will stop the action. If she doesn't, the AI will stay silent, and the accident will happen. The paper mathematically shows exactly how much "harm" is lost when this communication channel is broken and how long it takes to fix it through experience or better signaling.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →