← Latest papers
🤖 machine learning

Epistemology gives a Future to Complementarity in Human-AI Interactions

This paper addresses the theoretical and empirical limitations of human-AI complementarity by reframing it through an epistemological lens of computational reliabilism, arguing that it should serve as evidence of a reliable epistemic process to better guide decision-making and inform design and governance practices.

Original authors: Andrea Ferrario, Alessandro Facchini, Juan M. Durán

Published 2026-04-23
📖 6 min read🧠 Deep dive

Original authors: Andrea Ferrario, Alessandro Facchini, Juan M. Durán

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Why "Working Together" Isn't Always Enough

Imagine you are trying to solve a difficult puzzle. You have a friend (a human) and a super-fast robot (an AI).

For a long time, researchers have been obsessed with a concept called Complementarity. The idea is simple: If the human and the robot work together, they should solve the puzzle better than the human alone OR the robot alone.

The paper argues that while this sounds great, treating "Complementarity" as the only gold standard for success is a mistake. It's like judging a restaurant solely by how much food is on the plate, ignoring whether the food is fresh, safe to eat, or if the chef is actually skilled.

The authors, Andrea Ferrario and colleagues, propose a new way to look at this. They suggest we stop asking, "Did they beat the score?" and start asking, "Is this a reliable way to make decisions?" They use a philosophical framework called Computational Reliabilism to explain why.


The Problem: The "Complementarity Trap"

The paper identifies four big problems with how we currently measure human-AI teamwork:

  1. It's a Hindsight Mirror: You can only know if the team was "complementary" after you know the right answer. But in real life (like a doctor diagnosing a patient), you don't know the answer yet. You need to know now if you can trust the team.
  2. It Ignores the "Cost": What if the human and AI team gets a slightly better score, but it takes the human 10 hours of extra work to get there? That's a bad deal. The current definition doesn't care about the time or effort wasted.
  3. It's Too Narrow: A team could be "complementary" (better than either alone) but still be unfair, biased, or legally risky. Just because you improved the score doesn't mean you did the right thing.
  4. It's Hard to Achieve: In real-world studies, human-AI teams often don't beat the best of the two working alone. They often just make the same mistakes or get confused.

The Solution: The "Reliability Report Card"

Instead of a simple "Pass/Fail" on whether the team beat the solo scores, the authors suggest we view the Human-AI team as a Reliable Process.

Think of it like a Flight Safety Inspection.

  • The Old Way (Complementarity): "Did this plane fly 10% faster than a car?" (Irrelevant if the plane crashes).
  • The New Way (Reliabilism): "Is this plane safe to fly?" To answer that, we check three different things (called Reliability Indicators):

1. The Engine Check (Technical Performance)

This is where "Complementarity" fits in. Did the team actually perform better than the parts alone?

  • Analogy: If the pilot and the autopilot work together, do they land the plane smoother than the pilot alone or the autopilot alone?
  • The Catch: This is just one piece of evidence. It helps, but it doesn't guarantee safety.

2. The Blueprint Check (Epistemic Standards)

Is the task even being done correctly?

  • Analogy: Are we using the right map? Is the "destination" we are aiming for actually the right place? Are we measuring the right things?
  • Example: If an AI helps a judge decide on bail, is the AI measuring "risk of re-offending" correctly, or is it just measuring "how poor the defendant is"? If the blueprint is wrong, the team is unreliable, no matter how well they work together.

3. The Crew & Rules Check (Socio-Technical Practices)

Are the humans trained? Are there rules for when to panic?

  • Analogy: Is the pilot certified? Is there a clear rule for what to do if the engine sputters? Is there a logbook?
  • Example: If a doctor uses an AI, do they have training on when to ignore the AI? Is there a system to report if the AI starts making weird mistakes?

The "Efficiency" Meter: Is the Gain Worth the Pain?

The paper introduces a new concept: Efficient Complementarity.

Imagine you are baking a cake.

  • Scenario A: You and a robot mix the batter. You get a cake that is 1% tastier than if you baked it alone. But, it took you 5 extra hours to argue with the robot about the flour.
  • Scenario B: You and a robot mix the batter. You get a cake that is 20% tastier, and it took you 5 minutes.

The old way of thinking just says, "Scenario A is complementary because you did better!"
The new way says, "Scenario A is inefficient. The cost (5 hours of arguing) outweighs the tiny gain (1% tastier). Scenario B is efficient."

The authors argue we need to measure the Net Gain: How much better is the result, minus how much effort it took to get there?

Real-World Examples from the Paper

  1. The Dermatologist (The Good Team): A skin doctor uses an AI to spot cancer. They have training, the AI is well-tested, and they know when to double-check the AI. Sometimes they beat the AI alone, sometimes they beat the doctor alone. This is a Reliable Process because all three checks (Engine, Blueprint, Crew) are strong.
  2. The Student (The "Almost" Team): A student uses a free AI chatbot to do math homework. The AI makes mistakes, but the student catches them. Sometimes they get the right answer together. This is Partially Reliable. The "Engine" (complementarity) works sometimes, but the "Blueprint" (the AI isn't a certified math tool) and "Crew" (no teacher oversight) are weak. It's risky to trust this for a final exam.
  3. The Forensic Lab (The "No Gain" Team): A lab uses a voice-recognition AI. The human expert is already so good that the AI adds almost nothing. But, checking the AI takes time and slows down the court case. Here, Complementarity is low, but the process is still Reliable because the human is the expert, and the rules (Blueprint/Crew) are strict. Forcing them to "optimize for complementarity" would just waste time and money.

The Takeaway: What Should We Do?

The authors suggest we stop treating "Complementarity" as the only goal. Instead, we should use a Checklist for Trust.

When a hospital, a bank, or a government wants to use Human-AI teams, they shouldn't just ask, "Does it work better?" They should ask:

  • What is the protocol? (How do they talk to each other?)
  • What is the cost? (How much time/money does it take?)
  • Is it fair? (Does it hurt any specific group?)
  • Is it monitored? (What happens if the AI breaks?)

In short: Don't just look for the magic moment where Human + AI > Human + AI. Look for a reliable system where the human and AI work together safely, fairly, and efficiently, even if they don't always win the "scoreboard" game.

The paper ends by providing a Checklist (Table 1 in the appendix) that designers and regulators can use to ensure these teams are actually trustworthy, not just statistically "better."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →