← Latest papers
💬 NLP

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

This paper introduces a two-stage, multi-agent LLM framework that discovers interpretable surgical feedback quality criteria to automatically score live operating room interactions, demonstrating superior performance over prior methods in predicting feedback effectiveness and trainee behavioral adjustments.

Original authors: Rafal Kocielnik, J. Everett Knudsen, Steven Y. Cen, Jasmine Lin, Cherine H. Yang, Atharva Deo, Ujjwal Pasupulety, Peter Wager, Anima Anandkumar, Andrew J. Hung

Published 2026-05-26
📖 4 min read☕ Coffee break read

Original authors: Rafal Kocielnik, J. Everett Knudsen, Steven Y. Cen, Jasmine Lin, Cherine H. Yang, Atharva Deo, Ujjwal Pasupulety, Peter Wager, Anima Anandkumar, Andrew J. Hung

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a surgical operating room as a high-stakes orchestra. The surgeon in training is the lead violinist, and the attending surgeon (the teacher) is the conductor. For the music to be perfect, the conductor needs to give feedback. But here's the problem: sometimes the conductor's feedback is a gentle whisper, sometimes a sharp shout, and sometimes it's just a confusing murmur.

For years, researchers have tried to grade these musical directions by listening to what the conductor says (e.g., "fix the note" vs. "play louder"). But this new paper argues that how the feedback is delivered matters just as much as the words themselves. Did the teacher sound urgent? Was the instruction clear? Did they sound encouraging?

The authors, a team from Caltech and Cedars-Sinai, built a "smart robot listener" (a Large Language Model framework) to solve this. Here is how they did it, broken down into simple steps:

1. The "Brainstorming Party" (Discovery)

Instead of guessing what makes good feedback, the researchers asked a team of AI agents (think of them as five different expert music critics) to listen to thousands of real surgical conversations.

  • The Setup: They gave these AI agents a list of what "good" looks like (like "the student fixed their mistake") and asked them to figure out why some feedback worked and some didn't.
  • The Result: The AI agents came up with a long, messy list of ideas. The researchers then had a "consolidation meeting" where they grouped similar ideas together and distilled them down to six golden rules for good feedback:
    1. Encouraging: Does it boost confidence? (e.g., "Great job!")
    2. Urgent: Does it scream "Stop right now!"? (e.g., "Don't coagulate!")
    3. Actionable: Is it a specific step? (e.g., "Move the needle 2cm left.")
    4. Timely: Was it said while the action was happening, or hours later?
    5. Clear: Is it easy to understand, or confusing?
    6. Reflective: Does it make the student think? (e.g., "Why did you choose that?")

2. The "Robot Judge" (Scoring)

Once they had these six rules, they turned the AI into a judge. They fed it 4,200 real surgical feedback snippets and asked it to score each one from 1 to 5 on all six rules.

  • The Analogy: Imagine a teacher grading a student's essay. Instead of just giving a letter grade, this robot gives a detailed report card: "Your urgency was a 5/5, but your clarity was only a 2/5."

3. The Proof (Does it work?)

The team tested if these robot scores actually predicted what happened next in the operating room.

  • The Test: They looked at whether the student changed their behavior (like moving a tool correctly) or simply said "Okay" to acknowledge the teacher.
  • The Result: The AI's six rules were better at predicting student reactions than previous methods that just looked at the topic of the conversation.
    • Example: They found that "Urgent" and "Actionable" feedback made students move faster. But "Reflective" and "Clear" feedback made students talk back more.
    • Surprise: "Encouraging" feedback actually made students less likely to change their behavior or speak up. The paper suggests this is because encouragement is often used when the student is already doing a good job, so no change is needed!

4. The Human Check (Reliability)

To make sure the robot wasn't just hallucinating, they asked real human surgeons to grade the same feedback using the AI's six rules.

  • The Match: The humans and the AI agreed with each other quite well (about 60-70% agreement). This proves the rules are clear enough that both humans and machines can understand them.

Why This Matters (According to the Paper)

This framework is like giving every surgical teacher a "quality control dashboard."

  • It doesn't just tell you what was taught; it tells you how well it was taught.
  • It can automatically score thousands of hours of surgery without needing a human to sit there and take notes for days.
  • It helps identify if a teacher is being too vague, too late, or too confusing, which could help improve how surgeons are trained.

In short: The paper built an AI system that learned to grade surgical teachers not on their vocabulary, but on their delivery style. It found that being clear, urgent, and actionable is the secret sauce for getting students to actually fix their mistakes in real-time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →