← Latest papers
🩺 gastroenterology

ChatIBD: design, safeguards, and early international use of a guideline-grounded generative AI tool for inflammatory bowel disease (IBD) professionals

This paper describes the design, operational safeguards, and early international deployment of ChatIBD, a retrieval-augmented generative AI tool for inflammatory bowel disease professionals that demonstrated feasibility and global uptake over its first six months while emphasizing the need for further formal validation of its accuracy and clinical effectiveness.

Original authors: Chuah, C. S., Gros, B., Plevris, N.

Published 2026-07-30
📖 5 min read🧠 Deep dive

Original authors: Chuah, C. S., Gros, B., Plevris, N.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine the internet as a giant, chaotic library where every book is written by a different author, and some of those authors are notorious for making things up as they go along. This is the world of general-purpose Artificial Intelligence (AI). These digital brains are incredibly smart and can chat about almost anything, but they have a tricky habit called "hallucination." It's like a student who, when asked a specific question about a history event, confidently invents a story that sounds real but never actually happened. In the medical world, where a wrong fact could be dangerous, this is a big problem. Doctors need answers that are not just clever, but strictly true and backed by real, trusted sources.

To solve this, scientists are building "specialized librarians." Instead of letting the AI wander the whole library, they lock it inside a small, curated room containing only the most reliable medical textbooks and rulebooks. This technique is called Retrieval-Augmented Generation (RAG). Think of it as giving the AI a pair of glasses that only let it see the pages of the approved books. When you ask a question, the AI must find the answer inside those pages and show you exactly which page it found it on. If the answer isn't in the books, the AI is programmed to say, "I don't know," rather than making something up. This paper is about testing a new, specialized librarian designed specifically for doctors who treat Inflammatory Bowel Disease (IBD), a condition that causes painful swelling in the gut.


The Paper: ChatIBD's First Six Months

The researchers behind this study, a team of doctors who also happen to be the creators of a tool called ChatIBD, wanted to see how their new AI librarian performed in the real world. They didn't just test it in a lab; they let it loose on the internet for six months to see how doctors actually used it.

The Setup: A Guarded Chatbot
ChatIBD is a website where doctors can ask questions about IBD. But unlike a regular chatbot that might guess, this one is built with strict safety rules. It is grounded in a "curated corpus," which is a fancy way of saying a specific collection of 32 to 35 official medical guidelines from top organizations. When a doctor asks a question, the system doesn't just "think" of an answer; it first hunts through those specific guidelines to find the relevant text. It then uses an AI to summarize that text and, crucially, it must provide a clickable link to the exact page where the answer came from.

To make things even safer, the system has a special "dosing card" feature. If a doctor asks about how much medicine to give a patient, the AI doesn't guess the number. Instead, it pulls the dosage from a separate, manually checked database based on official European regulations and displays it in a clear card. The AI is only allowed to identify which drug is being asked about; the numbers themselves come from a rigid, unchangeable list.

The Experiment: Who Used It and How?
The study looked at data collected between October 1, 2025, and April 1, 2026. During this time, 913 people registered for the tool. Of those, 684 (about 75%) actually sent at least one message, and 349 (about 38%) were "active users," meaning they sent three or more messages.

The tool was used by doctors in 69 different countries and in 28 different languages. The most active users came from the United Kingdom and Spain. Interestingly, the tool was used mostly on weekdays (85.1% of the time), suggesting doctors were using it as a quick reference while working, rather than as a late-night study buddy.

What Did They Find?
The most common questions were about medications. Nearly half of all messages (46.6%) were about drug treatments. The system successfully triggered its special "dosing card" 1,332 times, showing that doctors were relying on it for quick, safe dosage checks.

The researchers also looked at what doctors were trying to do. The most common goal was "guideline synthesis," where a doctor asked the AI to summarize what the rules say about a specific topic. Other common requests included asking for definitions, checking safety warnings, or figuring out the order of treatments.

The Safety Check: Did It Slip Up?
The team kept a close eye on the tool's performance. They recorded 16 times when users gave explicit feedback. 15 of those were positive ratings. However, there was one negative rating that triggered a safety review. A doctor flagged a response about a drug called Risankizumab. The AI's sentence was slightly ambiguous and could have been misinterpreted to mean the drug was worse for certain patients than it actually was. The team reviewed this, realized the phrasing was risky, and changed the system to prevent that specific kind of confusion in the future.

What This Means (and What It Doesn't)
The study shows that ChatIBD is being used internationally and repeatedly by doctors who seem to find it useful for navigating complex medical rules. It suggests that a specialized, guideline-grounded AI can be a practical tool for professionals.

However, the authors are very careful not to overhype the results. They explicitly state that this study does not prove the tool is medically accurate, safe, or effective for treating patients. They didn't measure if patients got better or if doctors made fewer mistakes; they only measured how many people used the tool and how they used it. The single flagged error serves as a reminder that even with strict safeguards, human review is still necessary.

In short, ChatIBD passed its first real-world test by showing up, being used, and catching its own mistakes when pointed out. But the researchers emphasize that this is just the beginning. The tool is a promising "specialized librarian," but it still needs more rigorous testing to prove it can be fully trusted in the high-stakes world of patient care.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →