← Latest papers
💻 computer science

A generalizable foundation model for intraoperative understanding across surgical procedures

The paper introduces ZEN, a generalizable foundation model trained on over 4 million frames from 21 surgical procedures via self-supervised multi-teacher distillation, which consistently outperforms existing models across diverse downstream tasks and demonstrates robust cross-procedure generalization for intraoperative understanding.

Original authors: Kanggil Park, Yongjun Jeon, Soyoung Lim, Seonmin Park, Jongmin Shin, Jung Yong Kim, Sehyeon An, Jinsoo Rhu, Jongman Kim, Gyu-Seong Choi, Namkee Oh, Kyu-Hwan Jung

Published 2026-02-17
📖 5 min read🧠 Deep dive

Original authors: Kanggil Park, Yongjun Jeon, Soyoung Lim, Seonmin Park, Jongmin Shin, Jung Yong Kim, Sehyeon An, Jinsoo Rhu, Jongman Kim, Gyu-Seong Choi, Namkee Oh, Kyu-Hwan Jung

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to be a surgeon. You don't just want it to know how to hold a scalpel; you want it to understand the story of the surgery. Is the surgeon cutting a ligament? Are they stopping a bleed? Are they about to make a mistake?

For a long time, AI models for surgery have been like specialized interns. One intern is great at gallbladder surgery but gets lost in a heart surgery. Another is amazing at spotting tools but can't tell you what the surgeon is thinking or planning to do next. They are narrow, rigid, and require a massive amount of human teachers to learn even a single task.

This paper introduces ZEN, a new kind of AI that acts more like a seasoned Chief Resident who has seen thousands of surgeries and can adapt to almost any situation.

Here is the breakdown of how ZEN works, using simple analogies:

1. The Massive Library (The Data)

To become smart, ZEN didn't just read a few textbooks. The researchers fed it a massive library of over 4.3 million video frames from more than 4,700 surgeries.

  • The Analogy: Imagine a medical student who has watched every single surgery video ever recorded, covering 21 different types of procedures (like liver removal, gallbladder removal, and more). They didn't just watch; they studied the movements, the tools, and the flow of the operation.

2. The "Super-Teacher" Strategy (Multi-Teacher Distillation)

This is the secret sauce. Usually, you train an AI with one teacher. But the researchers realized that different teachers are good at different things.

  • Teacher A (MIS-DINOv2): This teacher is like a spatial genius. It's incredible at understanding where things are, how deep the cut is, and recognizing specific tools and anatomy. It sees the "3D map" of the surgery.
  • Teacher B (PeskaVLP): This teacher is like a language expert. It's great at connecting what it sees to words. It can look at a video and say, "This is a hemostasis step," or answer a question like, "What is the surgeon doing?"

The Innovation: Instead of choosing one teacher, they built ZEN to be a student who learns from both simultaneously. They used a technique called "distillation," which is like a master chef teaching an apprentice. The apprentice (ZEN) watches the two masters work and learns to combine the master chef's knife skills with the master's knowledge of recipes.

3. The "Zero-Shot" Superpower

Most AI models need to be retrained from scratch for every new hospital or every new type of surgery. If you show a standard AI a surgery it hasn't seen before, it often panics.

ZEN is different. It has generalization.

  • The Analogy: Imagine you teach a human how to drive a sedan. If you put them in a pickup truck or a sports car, they might be a little shaky, but they still know how to drive because they understand the principles of driving (steering, braking, looking ahead).
  • ZEN's Feat: The researchers tested ZEN on surgeries it had never seen during its training (like brain surgery or robotic procedures). Even without specific training on those, ZEN performed better than existing models. It understood the concept of surgery, not just the specific video it memorized.

4. What Can ZEN Actually Do?

The paper tested ZEN on 20 different "exams" (tasks), and it aced them all:

  • The Translator (Vision-Language): You can ask ZEN, "What is the surgeon doing right now?" or "How many tools are visible?" and it answers in plain English. It can even describe the next step in the surgery.
  • The Map Maker (Spatial Understanding): It can draw a perfect outline around every tool and organ in the video (segmentation) and even guess how deep the camera is looking (depth estimation).
  • The Time Traveler (Workflow): It can look at a video and say, "We are currently in the 'cutting' phase, and we are about to move to the 'stitching' phase."
  • The Safety Inspector (Skill Assessment): It can check if the surgeon has met the "Critical View of Safety" (a crucial safety checklist in gallbladder surgery) and tell you if it's safe to proceed.

Why Does This Matter?

Currently, surgical AI is like a calculator: it's great at one specific math problem but useless for writing a poem.
ZEN is like a Swiss Army Knife. It is a "Foundation Model," meaning it is a single, powerful brain that can be adapted for many different jobs:

  • Training: It can act as a real-time coach for medical students, pointing out mistakes or explaining what's happening.
  • Safety: It can alert a surgeon if they are about to cut the wrong thing.
  • Efficiency: It can automatically log what happened during a surgery, saving hours of paperwork.

The Bottom Line

The authors built ZEN to be the first "universal translator" for the operating room. By teaching it to learn from multiple expert teachers and feeding it a massive, diverse diet of surgical videos, they created an AI that doesn't just "see" the surgery—it understands it. This is a huge step toward making surgery safer, more consistent, and easier to teach.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →