The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
This paper introduces HEG-TKG, a system that utilizes hierarchical evidence-grounded temporal knowledge graphs to overcome the "Provenance Gap" in clinical AI by ensuring 100% verifiable citations and high resistance to errors for rare disease reasoning, addressing the tendency of frontier large language models to fabricate references.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Problem: The "Magic Trick" of AI Doctors
Imagine you go to a doctor, and they give you a diagnosis. But instead of saying, "I found this in your blood test," they say, "I just know this because I read a book once."
That is essentially what current "Frontier" Large Language Models (LLMs) do when they act as medical experts. They are incredibly smart and can often give the right answer. However, when you ask them, "Where did you get that information?" they often make up fake references.
The authors call this the "Provenance Gap."
- Provenance means "where something came from."
- The Gap is the space between the AI's answer and the proof that the answer is true.
The Analogy:
Think of a student taking a history exam.
- The AI writes a perfect essay.
- The Teacher asks, "Show me your sources."
- The AI hands over a list of books that look real but don't actually exist, or books that are about baking instead of history.
- The Result: The teacher can't trust the essay because they can't verify the facts. In medicine, this is dangerous. If a doctor trusts a fake citation, a patient could get the wrong treatment.
The Experiment: Testing the "Magic"
The researchers tested five of the smartest AI models on three pairs of rare, confusing diseases (like distinguishing between two types of muscle weakness that look identical at first).
- The Test: They asked the AI to diagnose patients and provide citations (proof) from medical journals.
- The Result:
- When not asked to cite, the AIs gave zero real proof.
- When forced to cite, the best AI only got 15% of the citations right. The rest were "hallucinations"—real-sounding numbers pointing to the wrong topics or fake papers.
- The Shock: Even if you ask the AI to "be honest," it still lies about the sources because it's trying to guess what a citation looks like, not actually finding one.
The Solution: HEG-TKG (The "Annotated Map")
To fix this, the team built a new system called HEG-TKG. Instead of letting the AI guess, they built a massive, structured "map" of medical knowledge before the AI ever starts talking.
The Analogy: The Librarian vs. The Improviser
- Old AI (The Improviser): It's like a brilliant actor on stage who makes up a story on the spot. It sounds convincing, but you can't check the script.
- HEG-TKG (The Annotated Librarian): Imagine a librarian who has a giant, organized filing cabinet. Every single fact in the cabinet is glued to a specific page number in a real medical journal.
- When you ask a question, the librarian doesn't guess. They pull the exact file, read the page, and hand you the page number.
- The "Time" Feature: This system is special because it doesn't just know what happens; it knows when it happens. It tracks the timeline of a disease (e.g., "Symptom A happens at age 3, Symptom B happens at age 10"). Most other medical maps are static; this one is a moving timeline.
How It Works (The Three-Layer Cake)
The system organizes evidence into three "tiers" of trust, like a quality control checklist:
- Gold Tier: Facts confirmed by top medical guidelines and multiple sources. (High confidence).
- Silver Tier: Facts found in multiple different studies. (Good confidence).
- Bronze Tier: Facts found in just one study. (Needs a double-check).
When the AI generates an answer, it pulls from this "Gold/Silver/Bronze" map. Every claim it makes is tied to a real, verifiable ID number (a PMID) that a human doctor can look up in seconds.
The Results: Trustworthy vs. Trusting
The researchers compared three groups:
- Vanilla AI: Just the raw AI guessing.
- Guideline-RAG: The AI reading raw text documents (like a PDF) without a map.
- HEG-TKG: The AI using the structured, time-aware map.
The Findings:
- Accuracy: All three groups were about equally good at giving the right medical advice.
- Trust: The "Vanilla" and "Raw Text" groups gave zero verifiable proof.
- HEG-TKG: Achieved 100% verifiability. Every single claim could be traced back to a real paper.
- The "Fake" Test: The researchers tried to trick the system by feeding it fake medical errors. The system caught 100% of the errors because the "trace" (the citation) showed the error immediately.
Why This Matters for You
- Safety: Doctors can't use AI if they can't check the work. This system lets a doctor look at an AI's answer and say, "Okay, I see the source. I can verify this."
- Rare Diseases: For rare conditions, a doctor might have never seen a case before. This system acts like a super-powered, verified encyclopedia that tells them exactly what to look for and when to look for it.
- Privacy: The system can run on a hospital's own computer (on-premise), meaning patient data never has to leave the building to be processed by a cloud AI.
The Bottom Line
The paper argues that accuracy isn't enough. An AI can be right but still be useless if you can't prove why it's right.
The "Provenance Gap" is the missing link in medical AI. By building a system that forces the AI to show its work using a structured, time-aware map of real medical facts, the researchers have created a tool that doctors can actually trust. It turns the AI from a "magic trick" into a "verified research assistant."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.