← Latest papers
💻 computer science

Measuring Successful Cooperation in Human-AI Teamwork: Development and Validation of the Perceived Cooperativity and Teaming Perception Scales

This paper introduces and validates two new scales, the Perceived Cooperativity Scale (PCS) and the Teaming Perception Scale (TPS), which effectively measure the subjective quality of human-AI teamwork across diverse contexts through three empirical studies.

Original authors: Christiane Attig, Christiane Wiebel-Herboth, Patricia Wollstadt, Tim Schrills, Mourad Zoubir, Thomas Franke

Published 2026-04-28
📖 6 min read🧠 Deep dive

Original authors: Christiane Attig, Christiane Wiebel-Herboth, Patricia Wollstadt, Tim Schrills, Mourad Zoubir, Thomas Franke

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to bake a cake with a new, high-tech robot assistant. Sometimes, the robot is amazing: it knows exactly when to mix, it tells you if the oven is too hot, and you feel like you're a perfect baking team. Other times, the robot is confusing: it burns the batter, ignores your instructions, or acts like it has no idea what it's doing.

This paper is about creating two special "report cards" (called scales) to measure exactly how good that baking partnership feels, from your perspective. The researchers, working with Honda and the University of Lübeck, wanted to move beyond just asking, "Do you like this robot?" to asking, "How well did we actually work together?"

Here is a breakdown of their work using simple analogies:

The Two Report Cards

The researchers realized that working with AI happens in two different ways, so they built two different measuring tools:

1. The "Single Moment" Report Card (Perceived Cooperativity Scale - PCS)

  • What it measures: How well the AI acted during one specific task.
  • The Analogy: Think of this like judging a single dance move. Did your partner lead clearly? Did they follow your lead? Did they stay in rhythm just for this one song?
  • The Theory: It's based on "Joint Activity Theory." It checks if the AI was transparent (you knew what it was thinking), reliable (it did what it said), and responsive (it listened to you).
  • The Result: They tested this on 409 people in three different scenarios: playing a cooperative card game, chatting with a text-bot, and using a car's trip planner. The card worked great! It could clearly tell the difference between a human partner (who got high scores), a smart-but-predictable rule-based robot (medium scores), and a robot that learned by trial and error and got confused (low scores).

2. The "Long-Term Team" Report Card (Teaming Perception Scale - TPS)

  • What it measures: The feeling that you and the AI have become a genuine team over time.
  • The Analogy: This isn't just about one dance move; it's about whether you feel like you are part of a dance troupe. Do you feel like you are both investing effort? Do you feel like you share a common goal? Do you feel like the robot is a "teammate" rather than just a tool?
  • The Theory: It's based on "Evolutionary Cooperation Theory." It looks at whether both sides are paying a "cost" (effort) to help the other succeed.
  • The Result: This scale has three parts:
    • The Team: "We are a unit."
    • The Partner: "You are helping me."
    • The Self: "I am helping you."
    • The Twist: The "Self" part was tricky. People tended to give themselves high scores no matter how bad the robot was. It's like a student who thinks, "I studied hard!" even if they didn't actually study. The researchers found this part of the scale is sensitive to the situation; it works better when people have more freedom to choose how hard they work.

How They Tested It (The Three Studies)

To make sure their report cards were accurate, they used three different "training grounds":

  1. The Card Game (Hanabi): Imagine a game where you can't see your own cards and must rely on hints from your partner.

    • The Test: People played with a human, a robot that followed strict rules, and a robot that learned on its own.
    • The Finding: The report cards successfully ranked them: Human > Rule-Based Robot > Learning Robot. The learning robot was often confusing, so people didn't feel like they were cooperating well.
  2. The Chatbot (LLM): People used a text-bot to fact-check articles.

    • The Test: They used bots that were helpful, bots that were helpful but added unwanted changes, and bots that failed.
    • The Finding: The scales picked up on the differences in how "cooperative" the bots felt.
  3. The Car Trip Planner: People used a Tesla trip planner to find charging stations.

    • The Test: This was a "boundary case." The car isn't really a "teammate"; it's just a tool.
    • The Finding: The "Single Moment" card (PCS) still worked fine here. However, the "Long-Term Team" card (TPS) wasn't used because a car trip planner doesn't really have a "mind" or "goals" to team up with. You can't really feel like a team with a GPS.

What They Learned (The Takeaways)

  • The "Single Moment" card is rock solid. It is a very reliable way to measure how well an AI acts during a specific task. It works for everything from chatbots to game bots.
  • The "Team" card is powerful but needs care. It measures the deep feeling of partnership. However, the part where people rate their own contribution is a bit biased. People tend to think they are great teammates even when the AI is terrible. This suggests that in rigid situations (like a card game where you must cooperate), people's self-ratings don't change much.
  • Context matters. If you are just using a tool (like a calculator or a GPS), the "Team" feeling doesn't really apply. But if you are working on a complex project where you and the AI need to rely on each other, these scales tell you if that partnership is working.

Why This Matters

Before this, we mostly had tools to ask, "Do you trust AI?" or "Is this AI smart?" This paper gives us a way to ask, "How well are we working together right now?" and "Do we feel like a team?"

This helps designers build better AI. If a robot gets low scores on "predictability," the designers know to make the robot's actions clearer. If the "Team" score is low, they know the robot needs to show more support and shared goals.

In short: The researchers built two rulers. One measures how well a robot dances in a single song. The other measures if you feel like you and the robot are in a dance troupe together. They tested these rulers on 409 people and found they work well, though the "dance troupe" ruler needs to be used carefully depending on how much freedom the human has to move.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →