← Latest papers
💻 computer science

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

SurgMotion is a video-native foundation model that leverages latent motion prediction and a massive 15M-hour dataset to overcome the limitations of pixel-level reconstruction, achieving state-of-the-art performance across diverse surgical video understanding tasks including workflow recognition, action detection, and skill assessment.

Original authors: Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Me
Published 2026-03-25
📖 4 min read☕ Coffee break read

Original authors: Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to be a surgeon. You show it thousands of hours of surgery videos.

The Old Way (The Problem):
Previous AI models were like a student who obsessively memorized every single pixel in the video. They spent hours trying to remember exactly how the smoke from a cautery tool looked, how the light reflected off a wet organ, or how the fluid moved. They got so bogged down in these tiny, messy details that they missed the big picture: What is the surgeon actually doing? They were like someone trying to understand a movie by staring at the grain of the film stock instead of the plot.

The New Way (SurgMotion):
The researchers behind SurgMotion realized that surgeons don't care about the smoke; they care about the movement of the tools and the relationship between the instrument and the tissue.

They built a new AI that learns differently. Instead of trying to redraw the picture pixel-by-pixel, it learns to predict what happens next based on the motion. It's like watching a dance and learning the steps and the rhythm, rather than trying to memorize the color of the dancer's shoes.

The Three Secret Ingredients

To make this work, the team added three special "training wheels" to the AI:

  1. The "Motion Spotlight" (Motion-Guided Prediction):
    Imagine a spotlight that only shines on the parts of the video that are moving. If a surgeon is cutting tissue, the spotlight gets brighter. If the camera is just hovering over a still organ, the spotlight dims. This forces the AI to ignore the boring, static background and focus entirely on the action. It teaches the AI to say, "Hey, look at the scissors moving, not the smoke!"

  2. The "Relational Mirror" (Spatiotemporal Self-Distillation):
    Surgery is a dance of relationships. The tool must touch the tissue, which must move away. The AI uses a "teacher-student" system. The "teacher" (a more stable version of the AI) shows the "student" how to keep these relationships consistent over time. It's like a dance instructor telling the student, "Don't just move your arm; make sure your hand stays connected to the partner's hand." This stops the AI from getting confused when the camera angle changes.

  3. The "Variety Enforcer" (Feature Diversity):
    Surgery videos can be very boring visually—lots of pink skin and smooth organs. If the AI isn't careful, it might get lazy and just say, "Everything looks the same." The researchers added a rule that forces the AI to find differences even in similar-looking scenes. It's like a teacher telling a student, "Even though these two apples look red, find the tiny difference in their shape." This keeps the AI sharp and ready for anything.

The Massive Library (SurgMotion-15M)

You can't learn to be a surgeon by watching only one type of surgery. The team created SurgMotion-15M, the largest library of surgical videos ever assembled.

  • The Scale: It's like a massive movie theater with 3,658 hours of footage.
  • The Variety: It covers 13 different body parts (from the brain to the colon) and 100+ different procedures.
  • The Goal: By feeding the AI this massive, diverse diet, it learns a "universal" understanding of surgery. It doesn't just know how to remove a gallbladder; it understands the concept of surgery, so it can handle a brain operation or an eye surgery just as well.

The Results: Why It Matters

When they tested this new AI, it didn't just win; it dominated.

  • Workflow Recognition: It can tell you exactly what stage of the surgery is happening (e.g., "We are now cutting the duct") much better than any previous model.
  • Action Triplet Recognition: It can break down complex scenes into "Tool + Action + Target" (e.g., "Scissors + Cutting + Artery") with incredible accuracy.
  • Skill Assessment: It can look at a surgeon's movements and tell you if they are a novice or an expert, judging the quality of the movement, not just the result.
  • Seeing in 3D: Even though it wasn't explicitly taught geometry, it learned to estimate depth (how far away things are) just by watching the motion, which is crucial for 3D reconstruction.

The Bottom Line

SurgMotion is a breakthrough because it stopped trying to be a camera and started trying to be a surgeon. By focusing on motion and relationships rather than pixels, it has created a universal "brain" for surgical AI that can adapt to almost any operation, any hospital, and any camera setup. This paves the way for AI that can act as a real-time assistant in the operating room, helping surgeons make better decisions and training the next generation of doctors.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →