DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
The paper introduces DigitalCoach, a multimodal dataset of human expert-novice computer coaching sessions, to reveal that current AI models struggle with effective teaching by relying on direct instructions rather than explanations and by failing to maintain visual grounding, thereby hindering deep learner engagement.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to learn how to use a complex piece of software, like a 3D design tool or a spreadsheet program. You have two options for getting help:
- The Human Coach: An expert sits next to you (or on a video call), watches your screen, and says, "I see you're stuck on that button. It's grayed out because you haven't selected the shape first. Let's try clicking the shape, then the button. Do you see why that happened?"
- The AI Coach: A smart computer program reads your screen and says, "Click the button. Then click the confirm box. Then save the file."
The paper DIGITALCOACH asks a simple but important question: Can the AI Coach teach you as well as the Human Coach?
To find out, the researchers built a massive "training gym" for AI. Here is the breakdown of what they did and what they found, using some everyday analogies.
1. The Dataset: Building the "Gym"
The researchers created a new dataset called DIGITALCOACH. Think of this as a giant library of recordings.
- The Content: They recorded 72 sessions where a real human expert taught a beginner how to use 5 different software programs (like Blender for 3D art, Excel for spreadsheets, and OnShape for engineering).
- The Detail: It wasn't just audio. They recorded the screen, every mouse click, every keystroke, and saved snapshots of the files. It's like having a high-definition replay of the entire learning process, not just the conversation.
- The Goal: They used these recordings to see how humans actually teach, and then tested if AI models could copy that style.
2. The Problem: The "Robot Teacher" vs. The "Human Mentor"
When the researchers tested the best AI models available today, they found a big gap between how robots teach and how humans teach.
The Human Coach (The "Mentor"):
- Explains the "Why": They don't just say "do this." They explain why you are doing it. "We are clicking this because it turns your flat shape into a 3D object."
- Checks for Understanding: They ask questions like, "Do you see where that button is?" or "What do you think happens next?"
- Adapts to the Screen: If you click the wrong thing, they notice immediately and say, "Oh, that button is disabled because..."
The AI Coach (The "Robot"):
- Just Gives Orders: The AI mostly just says, "Click here. Then click there." It's like a GPS that only gives turn-by-turn directions without ever explaining the map.
- Blind to the Screen: The AI often ignores what is actually happening on your screen. If you click the wrong button, the AI might keep telling you to click that same button, not realizing it's grayed out or broken.
- No Questions: It rarely asks if you understand. It just keeps pushing the next instruction.
3. The Experiment: Who Learned More?
The researchers put this to the test with real people.
- Group A was taught by a human expert.
- Group B was taught by an AI coach.
The Results:
- Group A (Humans): The learners understood the concepts better. They finished more of the tasks and, more importantly, they could do similar tasks on their own afterward. They learned the skill, not just the steps.
- Group B (AI): The learners got stuck more often. They followed the instructions blindly but didn't understand what they were doing. When the task got slightly different, they couldn't figure it out. They were like a passenger following a driver who refuses to explain the route; if the driver stops, the passenger is lost.
4. The Core Issue: "Grounding"
The paper uses a term called "Grounding."
- Imagine you are playing a video game with a friend. If your friend says, "Jump over the red block," but the block in front of you is blue, a "grounded" friend would say, "Wait, that's blue. Did you mean the red one behind you?"
- The Human Coach is grounded. They see the screen and match their words to what they see.
- The AI Coach is "ungrounded." It often forgets to look at the screen. It relies too much on the text conversation and not enough on the visual reality. It's like a teacher reading from a script while the student is actually doing something different.
5. The Conclusion
The paper concludes that while AI is getting very good at doing tasks for us (like writing code or moving files automatically), it is currently bad at teaching humans how to do those tasks themselves.
The AI acts more like a remote control that executes commands, rather than a coach who helps you build your own skills. To make AI coaches truly helpful, they need to learn to:
- Look at the screen in real-time.
- Ask questions to check if you understand.
- Explain the "why" behind the steps, not just the "what."
Until AI can do these things, it's better to think of it as a tool that does the work for you, rather than a teacher that helps you learn the work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.