Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy
This paper introduces EndoASR, a domain-adapted automatic speech recognition system that significantly improves transcription accuracy and medical term recognition for real-time human-AI collaboration in gastrointestinal endoscopy, demonstrating robust generalization and efficient edge deployment across multi-center clinical settings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a surgeon performing a delicate procedure inside a patient's body using a long, flexible camera (an endoscope). Your hands are busy holding the camera and tools, so you can't type or write notes. Instead, you talk out loud, describing what you see: "I see a small red bump on the left side of the colon, looks like a polyp, let's remove it."
In a perfect world, a computer listens to you and instantly writes a perfect medical report. But in the real world, this is like trying to have a serious conversation with a friend who only speaks a different language, while a loud construction crew is drilling next to you.
This paper introduces EndoASR, a new "smart listener" designed specifically to solve this problem for gastroenterologists. Here is the story of how they built it, explained simply:
1. The Problem: The "Generalist" vs. The "Specialist"
Think of standard voice assistants (like Siri, Alexa, or the big AI models everyone uses) as generalist librarians. They are great at reading books about history, cooking, or pop culture. But if you ask them to read a complex medical textbook, they stumble.
- The Vocabulary Gap: Doctors use very specific words like "Paris classification IIa" or "Boston Bowel Preparation Scale." A generalist librarian has never heard these words and guesses wrong.
- The Noise Gap: An endoscopy room is loud. There are beeping machines, suction hoses, and multiple people talking. A generalist librarian gets confused by the background noise.
The researchers tested these "generalist" librarians in a real endoscopy room, and they failed miserably, making too many mistakes to be useful.
2. The Solution: Training a "Specialist" Intern
Since they couldn't find enough real recordings of doctors talking (because of patient privacy rules), the researchers had to get creative. They built a synthetic training camp.
- Step 1: The Script (Language Adaptation): They took thousands of real, written medical reports and used a Text-to-Speech robot to read them aloud. This created a massive library of "fake" doctor voices speaking perfect medical terms. They taught their AI model to listen to this, turning it into a medical vocabulary expert.
- Step 2: The Noise Gym (Robustness): Just knowing the words isn't enough; the AI needs to hear them over the noise. They took that "fake" speech and mixed it with recordings of actual endoscopy room noise (beeping, suction, etc.). They trained the AI to ignore the chaos and focus on the doctor's voice.
This two-step process created EndoASR, a model that is both a medical dictionary and a noise-canceling headphone in one.
3. The Test Drive: From the Lab to the Real World
The researchers didn't just test this in a quiet lab. They took it on a road trip across five different hospitals.
- The Result: In these real-world tests, EndoASR was like a seasoned local guide compared to the confused tourist (the old models). It reduced errors by about 30% and got the medical terms right nearly 85% of the time (up from about 60% for the old models).
- Speed: It was incredibly fast. While other models took a moment to "think" (like a slow translator), EndoASR was almost instantaneous, allowing the doctor to speak and see the text appear on the screen in real-time without waiting.
4. Why It Matters: The "Teamwork" Effect
The most important part isn't just that the computer hears the words correctly; it's what happens after.
Imagine the AI is a team member helping the doctor.
- If the AI hears "polyp" correctly, it can instantly pull up the right treatment guide or fill out the insurance form.
- If the AI mishears "polyp" as "puppy," the whole team gets confused, the report is wrong, and patient safety is at risk.
By fixing the hearing (the ASR), the whole team (the doctor + the AI) works better together. The doctor can keep their hands on the tools and their eyes on the screen, trusting the AI to handle the paperwork.
The Bottom Line
This paper shows that to make AI useful in the messy, noisy, high-stakes world of medicine, you can't just use a "one-size-fits-all" tool. You have to build a specialist that speaks the doctor's language and wears noise-canceling headphones. EndoASR is that specialist, ready to be the reliable ears for the human-AI medical team.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.