← Latest papers
🤖 AI

AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

This paper investigates the distinct dynamics of delegation and adoption in human-AI question-answering teams, revealing that while collaboration outperforms individual agents, humans make suboptimal trust decisions due to confirmation bias and misaligned confidence, necessitating improved calibration and explanation mechanisms.

Original authors: Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig, Irene Ying, Tianyi Zhou, Jordan Boyd-Graber

Published 2026-05-28
📖 6 min read🧠 Deep dive

Original authors: Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig, Irene Ying, Tianyi Zhou, Jordan Boyd-Graber

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a high-stakes trivia night where you aren't just playing against other people, but you are paired up with a team of robot teammates. This is exactly what researchers did to figure out how humans and AI should work together. They didn't just ask people, "Do you trust the robot?" They put them in a real game where the robot could make mistakes, and the humans had to decide when to let the robot take the wheel and when to grab the steering wheel back.

Here is the breakdown of their findings using simple analogies:

The Game: A Trivia Tournament

The researchers set up a "Quiz Bowl" tournament.

  • The Players: 23 human trivia experts (people who have been playing for years) and 16 different AI robots (built by different programmers using various brainy models).
  • The Teams: Humans formed teams and had to "draft" two AI robots to be their teammates, kind of like picking players in a fantasy sports league.
  • The Two Ways to Play:
    1. The "Buzz" Phase (Delegation): Questions are read aloud, and anyone can buzz in to answer immediately. Here, the human has to decide: Do I let my robot teammate buzz in and answer on its own, or do I mute it and do it myself? This is like deciding whether to let your GPS drive the car while you nap, or if you should keep your hands on the wheel.
    2. The "Bonus" Phase (Adoption): If the team gets a question right, they get a bonus round with three parts. Here, the human guesses first, then sees what the robots think, along with their confidence scores and explanations. The human then decides: Do I stick with my guess, or do I switch to what the robot says? This is like checking a recipe, seeing what the AI suggests, and deciding if you trust its advice enough to change your dish.

What They Found: The "Trust Gap"

The team found that humans and AI are actually a great team together, but they aren't perfect at working together yet.

1. The "Too Much Trust" Problem (Over-reliance)
Sometimes, humans get the answer right, but the robot gives a confident, wrong answer. The human sees the robot's confidence and thinks, "Oh, the robot must know something I don't," and switches to the wrong answer.

  • The Analogy: It's like you know the capital of a country is Paris, but your GPS confidently says it's London. You panic and change your destination to London, even though you were right all along. This happened about 1.7% of the time.

2. The "Not Enough Trust" Problem (Under-reliance)
This was actually more common. Sometimes the human gets the answer wrong, and the robot gets it right. But the human ignores the robot's correct answer and sticks with their own mistake.

  • The Analogy: You think the answer is "The Moon," but your robot teammate says "The Sun" and explains exactly why. You ignore the explanation because you're stubborn or confident in your wrong answer. This happened about 3.9% of the time.
  • The "Echo Chamber" Effect: The study found that if a human is wrong, and one robot agrees with them, the human becomes super stubborn. They trust the robot's agreement so much that they refuse to change their mind, even if the other robot is right. This is called confirmation bias.

3. The Robot's "Fake Confidence"
The robots often gave confidence scores (like "I am 90% sure") that didn't match reality.

  • The Analogy: Imagine a weatherman who says, "I am 100% sure it will rain," but it's a sunny day. The humans in the study learned that the robot's confidence score wasn't a reliable compass. When the human and robot disagreed, the robot's confidence score was basically a coin flip.

How Humans Actually Make Decisions

The researchers looked at why humans trusted or ignored the robots. They found a funny mismatch:

  • What makes a robot actually correct? Deep reasoning, understanding the clues, and citing specific evidence from the question.
  • What makes a human trust a robot? Surface-level stuff. Humans loved it when the robot used quotes from the question or when the robot's answer sounded similar to their own thoughts.
  • The Takeaway: Humans were fooled by "style over substance." They trusted the robot that sounded fancy or used quotes, even if the robot was wrong. They ignored the robot that gave a deep, evidence-based explanation if it didn't look "familiar."

The Golden Rules for Better Teamwork

The paper suggests five simple rules to fix these issues, based on what they saw in the game:

  1. Let Humans Control the Switch: Don't just have an "On/Off" button for AI. Let humans mute the robot for specific topics (like "Don't let the robot answer music questions") so they can use their judgment on where the robot is weak.
  2. Standardize the Confidence: Robots need to stop lying about how sure they are. If a robot is wrong, it shouldn't say it's 90% sure.
  3. Show the Track Record: Instead of just showing the robot's answer, show the humans a history: "This robot is great at history but terrible at math." This helps humans know when to listen.
  4. Fix the "Under-Reliance": Since humans are more likely to ignore a correct robot than to blindly follow a wrong one, systems should highlight when the robot is confident on topics the human usually struggles with.
  5. Anchor Explanations in Evidence: This is the big one. When a robot explains its answer, it shouldn't just say "I think so." It should say, "I think so because the question mentioned [specific clue]." Humans are much more likely to trust an answer when they can see the "receipts" (the specific evidence) rather than just a confident-sounding paragraph.

The Bottom Line

Humans and AI are a powerful team, but they are currently driving with mismatched maps. Humans need to stop trusting the robot just because it sounds confident or uses quotes, and robots need to stop pretending to be sure when they aren't. If they can learn to trust each other based on evidence rather than confidence, they can solve problems neither could solve alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →