← Latest papers
💻 computer science

Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles

This paper introduces "Interaction Readiness," a framework that distinguishes between content accuracy and role-specific behavioral performance to help teams build and evaluate AI agents by defining interaction specifications and operationalizing key agent capabilities like authority calibration and breakdown repair.

Original authors: Sudhir Alladi Venkatesh

Published 2026-08-14
📖 5 min read🧠 Deep dive

Original authors: Sudhir Alladi Venkatesh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a room where a robot has been hired to be your personal tutor. You might think the only thing that matters is whether the robot knows the right answers. If it can recite the periodic table or solve a math problem without making a typo, you'd probably call it a success, right? But here's the twist: being a good tutor isn't just about having a giant brain full of facts. It's about knowing when to speak, how to speak, and what you are actually allowed to do in that specific moment. This paper lives in the world of Artificial Intelligence, specifically looking at how we build "agents"—AI programs designed to act like humans in specific jobs, like a coach, a doctor, or a teacher. The author is asking a question that most people haven't thought to ask: What happens when an AI is factually perfect but socially clumsy? They introduce a new idea called "Interaction Readiness," which is basically a checklist to see if an AI can actually be the role it's pretending to play, not just say the right words.

The paper argues that we currently have a blind spot. We are great at testing if an AI is safe, polite, and factually correct. But we are terrible at testing if it understands the "social rules" of its job. Think of it like a play. If an actor memorizes every line perfectly but stands in the wrong spot, talks to the wrong person, or laughs at a sad moment, the play fails—even if the words were correct. The author suggests that for AI to work in real roles, we need to stop just checking the "content" (the words) and start checking the "interaction" (the behavior). They propose that an AI needs to understand four things: the purpose of the conversation, the limits of its authority (what it's allowed to do), the right tone of voice, and how to fix things when the conversation goes off the rails.

To test this, the researchers looked at a real-world dataset called StudyChat, which contains 937 conversations between university students and an AI tutor. This AI was placed in a classroom setting but was only told to be a "helpful assistant." It wasn't given a specific rulebook on how to act like a tutor. The team hand-scored 50 of these conversations to see if the AI was "interaction ready."

Here is what they found, and it's a bit surprising. The AI was often factually correct and very fluent. It sounded smart. But it failed the "tutor" role in a very specific way. The biggest problem was authority calibration. The AI didn't know when to hold back. In one famous example, a student asked the AI to write a short reflective essay for them based on a few notes. The AI happily wrote the whole essay. Factually? Perfect. As a tutor? A disaster. A real tutor would know that writing the essay for the student defeats the purpose of the assignment. The AI was too "helpful." It treated every request like a command to just do the work, rather than a chance to help the student learn.

The paper also showed that being "good at talking" and "good at knowing facts" are two completely different things. They found cases where the AI gave the wrong technical advice but was so polite and well-structured that the student trusted it anyway. Conversely, they found cases where the AI gave the right answer but did it in a way that felt robotic or missed the student's confusion. The author calls this "Social Fit." The AI in the study was like a person who knows the map but doesn't know how to drive the car; it could tell you the destination, but it didn't know how to navigate the traffic of the conversation.

The researchers suggest that the solution isn't just to make AI smarter. It's to give it a better "job description." Instead of just saying "be helpful," engineers need to write an "interaction specification" that tells the AI exactly what its role is. For a tutor, that means: "You can explain concepts, but you cannot write the graded assignment for the student." "If the student is stuck, ask a question instead of giving the answer." "If the student is frustrated, slow down." The paper suggests that without these specific rules, AI will default to being a generic "helpful assistant," which works sometimes but fails often when the situation gets complicated.

In short, the paper suggests that we are building AI agents that are like excellent encyclopedias but terrible teammates. They can recite the facts, but they don't understand the game they are playing. The author isn't saying this is a solved problem or that AI is broken forever. They are just pointing out that we need a new way to test these robots. We need to check if they can read the room, respect the rules of their job, and know when to stop helping so the human can actually learn. If we don't add this layer of "Interaction Readiness" to our testing, we might end up with AI that is technically perfect but completely useless at the actual job it was hired to do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →