← Latest papers
🤖 AI

Hybrid-Code v2: Zero-Hallucination Clinical ICD-10 Coding via Neuro-Symbolic Verification and Automated Knowledge Base Expansion

Hybrid-Code v2 is a neuro-symbolic framework that achieves zero Type-I hallucination in clinical ICD-10 coding by integrating neural candidate generation with a symbolic verification layer and automated knowledge base expansion, thereby delivering high precision and coverage without the safety risks of purely neural approaches.

Original authors: Yunguo Yu

Published 2026-03-24
📖 4 min read☕ Coffee break read

Original authors: Yunguo Yu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a librarian trying to organize a massive, chaotic library of patient stories. Every story needs to be tagged with a specific, official code (like a barcode) so the hospital knows how to bill insurance and track diseases. This is called ICD-10 coding.

The problem is that there are over 70,000 possible codes, and the rules for using them are incredibly strict. If you tag a story with the wrong code, it's like putting a "Poison" label on a jar of honey. It can cause billing fraud, confuse doctors, or even endanger patients.

For a long time, we've had two ways to do this tagging:

  1. The "Human Rulebook" (Rule-Based Systems): Imagine a librarian who only uses a tiny, strict checklist. They will never make a mistake because they only pick codes they are 100% sure of. But because the checklist is so small, they miss 55% of the books. They can't tag complex stories.
  2. The "Super-Intelligent Robot" (Neural AI): Imagine a robot that has read every book in the library. It is amazing at finding the right tags for complex stories. But, the robot has a bad habit: sometimes, when it's not sure, it hallucinates. It invents a fake code that sounds real but doesn't exist (like a barcode for "Dragon Disease"). In a hospital, making up a code is dangerous.

Enter "Hybrid-Code v2": The Smart Librarian with a Safety Net.

This paper introduces a new system that combines the best of both worlds. Think of it as a two-step assembly line:

Step 1: The Creative Generator (The Neural Part)

First, a smart AI (like the robot) reads the patient's story and suggests a list of possible codes. It's fast, flexible, and good at understanding complex language.

  • The Risk: Sometimes, this AI might suggest a code that doesn't actually exist in the official library.

Step 2: The Strict Gatekeeper (The Symbolic Part)

Before any code is ever written down, it must pass through a "Gatekeeper." This isn't a brain; it's a rigid, unbreakable rulebook. The Gatekeeper checks the AI's suggestions against the official, real ICD-10 codebook.

  • The Check: "Does this code actually exist?" "Is the patient actually sick with this (not just 'ruled out' or 'history of')?" "Do these two codes fight each other?"
  • The Result: If the AI suggested a fake code, the Gatekeeper slams the door. Zero fake codes get through.

The Magic Trick: Teaching the Gatekeeper to Grow

The old problem with rulebooks was that they were too small and took years for humans to write.
Hybrid-Code v2 has a secret weapon: Automated Expansion.
Imagine the Gatekeeper has a robot assistant that reads thousands of new patient stories, finds patterns humans missed, and suggests new rules to add to the Gatekeeper's book.

  • A human expert then quickly double-checks these new rules.
  • If they are good, they get added.
  • Now the Gatekeeper knows more codes without needing a human to write every single rule from scratch.

The Results: The Perfect Balance

The researchers tested this system on 5,000 patient stories. Here is what happened:

  • The "Fake Code" Problem: The smart robots (AI models) made up fake codes 6% to 18% of the time. The Hybrid-Code system made zero fake codes. It was perfect at safety.
  • The "Missing Code" Problem: The old rulebooks missed half the stories. The Hybrid-Code system found the right codes for 85% of the stories, which is almost as good as the smart robots.
  • The "Accuracy" Problem: When it did pick a code, it was right 88% of the time.

Why This Matters

Think of it like a self-driving car.

  • Old AI: Drives fast and handles traffic well, but sometimes it hallucinates a stop sign that isn't there and crashes.
  • Old Rules: Drives very slowly and safely, but it can't handle complex intersections, so it gets stuck.
  • Hybrid-Code: It drives fast and handles complex traffic, but it has an unbreakable safety belt that physically prevents it from driving off a cliff.

In short: This paper proves you don't have to choose between a smart system and a safe system. By adding a "fact-checker" layer to an AI, you can get the speed and smarts of a robot with the safety guarantees of a strict rulebook. This is a huge step toward trusting AI in hospitals.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →