Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos
This paper introduces Colon-Bench, a comprehensive benchmark featuring over 500 full-procedure colonoscopy videos with dense, multi-modal annotations generated via an agentic workflow, which is used to rigorously evaluate and improve Multimodal Large Language Models for medical lesion detection and analysis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to be a world-class colonoscopy doctor. The problem is, colonoscopies are like looking for a needle in a haystack, but the haystack is moving, blurry, and covered in mud. Furthermore, the "needles" (lesions like polyps or ulcers) come in all different shapes, sizes, and colors.
Until now, the "textbooks" (datasets) we used to train these robots were too simple. They mostly showed pictures of just one type of polyp, like a flashcard with only apples. But real life has apples, oranges, bruises, and cracks in the skin.
Enter "Colon-Bench."
This paper introduces a massive new training ground for AI, created by researchers at KAUST. Here is the story of how they built it and what they discovered, explained simply.
1. The Problem: The "Needle in a Moving Haystack"
Colon cancer is a huge killer, but it's preventable if caught early. The only way to catch it is a colonoscopy, which is a long, uncomfortable video procedure.
- The Challenge: Doctors have to watch hours of video to find tiny, hidden problems.
- The AI Gap: We want AI to help, but to teach an AI, you need thousands of videos where every single problem is marked with a sticky note (annotation). Doing this manually is like trying to paint the entire ocean; it takes too long and is too expensive.
2. The Solution: The "Robot Factory" (Agentic Workflow)
The researchers didn't just hire more humans to draw boxes around lesions. Instead, they built a multi-stage robot assembly line (called an "Agentic Workflow"). Think of it like a high-tech quality control factory:
- Stage 1: The Scout (AI Proposal): A smart AI scans the video and says, "Hey, I think I see something suspicious here!" It marks thousands of potential spots.
- Stage 2: The Detective (Verification): Another AI checks the Scout's work. "Are you sure? Or is that just a shadow?" It filters out the fake alarms.
- Stage 3: The Tracker (Box & Mask): Once a spot is confirmed, a third AI draws a tight box around it and follows it as it moves through the video, even if it gets hidden behind stool or fluid.
- Stage 4: The Human Expert (The Final Boss): A real surgeon looks at the top candidates. Because the robots did the heavy lifting, the surgeon only has to say "Yes" or "No" to the best 40% of the clips.
The Result: They turned 60 hours of raw video into a massive, verified library called Colon-Bench.
- 528 videos of full procedures.
- 14 different types of "bad guys" (not just polyps, but bleeding, ulcers, parasites, etc.).
- 300,000+ bounding boxes and 213,000+ masks (digital outlines).
- 133,000 words of clinical descriptions (like a doctor's diary entry for every clip).
3. The Test Drive: Can AI "See" Like a Doctor?
They used this new library to test the smartest AI models in the world (like Gemini, GPT-5, and Qwen). They asked the AI three types of questions:
- The Quiz (VQA): "What is this lesion, and where is it?"
- The Hunt (Detection): "Find the polyp in this video."
- The Trace (Segmentation): "Draw the exact outline of the polyp as it moves."
The Big Surprise:
The researchers expected specialized medical AI to win. Instead, the general-purpose "Super-Brain" AIs (Multimodal Large Language Models) crushed it.
- Gemini 3 Pro and Gemini 3 Flash became the top doctors, accurately identifying and tracking lesions better than the specialized tools.
- In fact, these general AIs were so good at localization that they beat the famous "Segment Anything" model (SAM-3) by a huge margin. It's like a generalist detective solving a medical mystery better than a specialist because they have seen so much more of the world.
4. The Secret Weapon: "Colon-Skill"
Even the smartest AIs make mistakes. They might confuse a "sessile polyp" (flat) with a "pedunculated polyp" (on a stalk), or mistake a hole (diverticulum) for a tumor.
The researchers analyzed where the AIs failed and wrote a "Cheat Sheet" (which they call Colon-Skill).
- The Analogy: Imagine you are taking a driving test. You know the rules, but you keep failing at parallel parking. Someone hands you a note that says: "Remember: Turn the wheel early, check the mirror, and don't hit the curb."
- The Result: When they gave this "Colon-Skill" note to the AIs before they took the test, their scores jumped by up to 9.7%. They didn't need to be retrained; they just needed a better hint.
Why Does This Matter?
- For Patients: This is a step toward AI that can watch colonoscopy videos in real-time, alerting doctors to tiny cancers they might miss, making the procedure safer and more effective.
- For AI: It proves that general AI models, when given the right data and the right "hints," can handle incredibly complex, messy medical tasks.
- For the Future: They showed that we can build massive, high-quality medical datasets without hiring armies of humans, using a smart mix of robots and a few expert humans.
In a nutshell: The researchers built a massive, high-definition "training gym" for AI doctors. They found that general AI super-brains are surprisingly good at spotting colon issues, and with a little bit of "study guide" (the Colon-Skill), they can become even sharper. This could soon help save lives by making colon cancer screening cheaper, faster, and more accurate.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.