← Latest papers
🤖 machine learning

Towards Actionable Surgical Team Dynamics: from Teamwork to Counterfactual Annotations

This paper introduces a comprehensive, multimodal dataset of real operating room recordings that integrates standardized teamwork evaluations, multi-level interaction annotations, and novel counterfactual descriptions of alternative outcomes to enable advanced computational modeling of surgical team dynamics and the development of AI-assisted collaborative systems.

Original authors: Vincenzo Marco De Luca, Antonio Longa, Andrea Passerini

Published 2026-08-25
📖 1 min read☕ Coffee break read

Original authors: Vincenzo Marco De Luca, Antonio Longa, Andrea Passerini

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Towards Actionable Surgical Team Dynamics

Problem Statement
Effective teamwork in high-stakes environments like operating rooms (OR) is critical for patient safety, relying heavily on non-technical skills such as communication, coordination, and situational awareness. While recent advances in surgical data science have produced multimodal datasets (e.g., Cholec80, MM-OR), these resources primarily focus on visual scene understanding, procedural phase recognition, and action detection. They lack integrated representations of team-level phenomena, such as speaker-specific conversational structures, non-technical skill assessments, and the dynamics of interaction breakdowns. Furthermore, existing datasets are often fragmented across modalities and annotation schemes, and they rarely provide explanatory insights into why teamwork performance degrades. Current computational models often reduce team performance to a single score, overlooking the complex interplay of individual, interpersonal, and collective processes, and fail to provide interpretable accounts necessary for trust and intervention in clinical settings.

Methodology
The authors propose a comprehensive extension of the publicly available MM-OR dataset, which contains synchronized multimodal recordings of robotic knee replacement procedures. The methodology involves three primary technical components:

  1. Multimodal Enrichment Pipeline:

    • Speaker Diarization: Manual refinement of speaker turns to ensure high temporal accuracy in the noisy OR environment, enabling analysis of turn-taking, interruptions, and participation balance.
    • Bilingual Transcription: Creation of aligned German (original) and English transcripts to facilitate annotation review and cross-lingual modeling.
    • Data Structure: The resource maintains hardware-synchronized modalities including 5 RGB-D cameras, 3 RGB cameras, 3 microphones, robotic logs, and 3D tracking, all aligned at 1 FPS.
  2. Multi-Level Annotation Framework:
    The authors introduce a hierarchical annotation schema covering 169 six-minute video clips, annotated by three experts using established frameworks:

    • Team-Level: Utilizing OTAS (Observational Teamwork Assessment for Surgery) to rate five constructs (Communication, Coordination, Cooperation, Leadership, Monitoring) on a 0–6 Likert scale.
    • Interaction-Level: Adapting the HMT (Human-Machine Teaming) framework for human-only teams to assess pairwise interactions across five macro-categories (Communication, Joint Information Processing, Coordination, Interpersonal Relationship, Motivation).
    • Individual-Level:
      • Dynamic: Using NOTSS (Non-Technical Skills for Surgeons) to assess Situation Awareness, Decision Making, Communication/Teamwork, and Leadership.
      • Static/Stable: Integrating BFTQ (Big-Five-In-Teamwork), SYMLOG (leadership/interpersonal behavior), and GLIS (General Leadership Impression Scale) to capture stable traits and emergent leadership.
  3. Counterfactual Annotation Layer:
    A novel contribution where annotators identify specific events or behaviors that most negatively impacted teamwork within a segment. These annotations include:

    • Temporal Intervals: Short durations of the triggering event.
    • Attribution: Identification of the affected team construct and the responsible agent(s).
    • Quality Ratings: Five dimensions evaluated on a 1–5 ordinal scale: Criticality (impact on performance), Plausibility (realism of the alternative), Confidence (annotator certainty), Minimality (sufficiency of the change), and Expected Enhancement (potential performance gain).
    • Rationales: Textual explanations describing the event, consequences, and suggested alternative behaviors.

Key Contributions

  • Enriched Multimodal Dataset: The first public OR dataset to combine raw multimodal signals with hand-crafted speaker-aware and language-aware representations (diarization and bilingual transcripts).
  • Multi-Level Annotation Framework: A unified resource providing simultaneous team-level, interaction-level, and individual-level assessments derived from multiple validated theoretical frameworks (OTAS, HMT, NOTSS, BFTQ, SYMLOG, GLIS).
  • Counterfactual Annotation Protocol: A novel layer capturing explanatory signals for teamwork degradation, moving beyond descriptive scores to identify specific behavioral triggers and plausible alternative outcomes.
  • Benchmark Tasks: The paper establishes baseline performance for three distinct research directions:
    1. Feature Enrichment: Empirically assessing and quantifying the utility of speaker-aware and transcript-based representations, demonstrating they provide complementary information and improve teamwork prediction over raw audio/visual features.
    2. Multi-Level Prediction: Showing that temporal-relational architectures (specifically TE-ReNN) outperform static and purely temporal models in predicting constructs across all annotation levels.
    3. Counterfactual Detection: Proving that multimodal signals contain predictive patterns for identifying events associated with teamwork degradation, achieving performance significantly above chance.

Results

  • Annotation Distribution: The dataset exhibits a moderate class imbalance, with most ratings concentrated around intermediate-to-positive scores (e.g., 62.7% of clips rated 4 or 5 for Communication in OTAS), reflecting generally competent but non-exceptional team performance in the recorded procedures.
  • Counterfactual Insights: Approximately 70% of counterfactual events relate to communication, monitoring, or cooperation failures. Most events are attributed to single actors (mean 1.3 actors), with criticality and confidence scores generally low-to-moderate, suggesting annotators viewed these as moderate inefficiencies rather than catastrophic failures.
  • Model Performance:
    • Feature Impact: Models utilizing diarization and transcripts consistently outperformed those using only raw audio and computer vision features.
    • Architecture: TE-ReNN (Temporal-Relational Neural Networks) achieved the highest F1-macro scores across all tasks (e.g., 74.2% for OTAS, 76.7% for BFTQ), confirming that jointly modeling temporal evolution and interpersonal dependencies is essential for surgical teamwork analysis.
    • Counterfactual Detection: Despite the rarity of events, models achieved F1-macro scores well above random chance, validating the learnability of these explanatory signals.

Significance and Claims
The paper claims to shift the paradigm of surgical data science from purely descriptive workflow modeling to integrated, multimodal, and explainable modeling of team dynamics. By providing a unified resource that links raw sensor data to high-level social constructs and counterfactual explanations, the work enables:

  • The study of how individual actions and interaction patterns jointly contribute to team outcomes.
  • The development of AI-assisted collaborative systems capable of identifying teamwork breakdowns and suggesting improvements.
  • A foundation for predictive modeling and behavioral analysis in high-stakes domains.

The authors modestly acknowledge that the counterfactual annotations represent expert judgments of plausible alternatives rather than experimentally validated causal mechanisms. They also note that the reliance on manual annotation limits scalability, suggesting future work may explore semi-automated pipelines. Nevertheless, the resource is positioned as a critical step toward interpretable, human-centered AI for surgical teams.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →