← Latest papers
💬 NLP

Automated Coding of Communications in Collaborative Problem-solving Tasks Using ChatGPT

This study demonstrates that ChatGPT can effectively automate the coding of communication data for collaborative problem-solving assessments, though its performance varies by model, framework, and task, with newer reasoning-focused models not necessarily outperforming others and prompt refinement yielding inconsistent improvements.

Original authors: Jiangang Hao, Wenju Cui, Patrick Kyllonen, Emily Kerzabi, Lei Liu, Michael Flor

Published 2026-03-04
📖 5 min read🧠 Deep dive

Original authors: Jiangang Hao, Wenju Cui, Patrick Kyllonen, Emily Kerzabi, Lei Liu, Michael Flor

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to grade a class project where students work in pairs to solve a mystery. The students can't talk out loud; they have to type their thoughts into a chat box. To figure out how well they are collaborating, you have to read every single message and sort them into categories like "Sharing an Idea," "Arguing a Point," or "Just saying hello."

If you have 10 students, this takes a few hours. If you have 10,000 students, this takes years. This is the "bottleneck" the paper talks about: grading these chats by hand is too slow and expensive to be useful for big tests.

The researchers asked a simple question: Can we hire a super-smart AI (ChatGPT) to do the grading for us?

Here is the breakdown of their experiment, explained with some everyday analogies:

1. The Experiment: The AI Intern vs. The Human Expert

The researchers treated the AI like a new intern. They gave the AI a "rulebook" (a coding framework) and a pile of chat logs from five different team tasks (some were science puzzles, others were negotiation games). They asked the AI to sort the messages just like a human expert would.

They tested four different versions of the AI:

  • The Standard Worker: GPT-4 and GPT-4o (the reliable, all-around workers).
  • The "Thinker" Models: GPT-o1-mini and GPT-o3-mini (newer models designed to "think" harder before answering, like a student who pauses to solve a complex math problem).

The Big Surprise: The "Thinker" models didn't do better. In fact, the standard, reliable worker (GPT-4o) was the best at this specific job. It's like hiring a genius philosopher to organize a library; they might overthink the Dewey Decimal System, while a regular librarian just gets it done faster and more accurately.

2. The Results: When the AI Shines and Stumbles

The AI didn't get a perfect score every time, but it did surprisingly well in two specific scenarios:

  • The "General Knowledge" Tasks (The AI Wins): When the tasks were about general skills like negotiating a budget or picking an apartment, the AI was almost as good as the human experts. It could tell the difference between "sharing info" and "asking for help" just fine.
  • The "Science" Tasks (The AI Struggles): When the tasks involved complex science terms (like "condensation" or "volcanic seismic activity"), the AI got confused.
    • Analogy: Imagine asking a brilliant translator to translate a conversation between two doctors discussing heart surgery. Even if the translator knows English perfectly, they might miss the nuance of a specific medical term. The AI struggled with the "jargon" of the science tasks.

3. The Rulebook Matters More Than the Worker

The researchers found that the quality of the rulebook mattered more than the intelligence of the AI.

  • Rulebook A (Theoretical): This was a rulebook written by academics based on pure theory. It was abstract and hard to follow. The AI struggled with this.
  • Rulebook B (Practical): This rulebook was built using real data and examples. It was clear and concrete. The AI crushed this one.

Analogy: If you give a robot a rulebook that says "Be nice," it might get confused. But if you give it a rulebook that says "If someone says 'Hello', reply 'Hi'," it works perfectly. The AI needs clear, practical instructions, not abstract theories.

4. Can We Fix the AI by Showing It Its Mistakes?

The researchers tried a "correction" strategy. They took the messages the AI got wrong, showed them to the AI, and said, "Hey, you messed up here; here is the right answer. Try again."

  • Task 1 (Condensation): This didn't help. It was like trying to teach a student by showing them one wrong answer; they just got confused about the other questions.
  • Task 2 (Volcano): This did help! The AI improved its score. It's like giving a student a specific practice quiz on the topic they missed; they learned from it.

The Bottom Line: The AI is a "Co-Pilot," Not the Pilot

The paper concludes that ChatGPT is a fantastic tool for speeding up the process of grading collaborative chats, but it's not ready to replace humans entirely.

  • Think of it this way: If you are grading 1,000 essays, you don't want to read every single word yourself. You let the AI read them first and sort them into "Good," "Okay," and "Needs Work." Then, a human expert just double-checks the tricky ones.
  • The Catch: You have to be careful with how you ask the AI to do it. You need a clear, practical rulebook, and you shouldn't assume the newest, most expensive AI model will automatically be the best.

In short: We can now use AI to help us understand how people work together, which means we can finally test "21st-century skills" on a massive scale without spending a fortune. But we still need human experts to hold the steering wheel.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →