← Latest papers
💬 NLP

Human-AI Co-design for Clinical Prediction Models

The paper introduces HACHI, an iterative human-in-the-loop framework that leverages AI agents to rapidly explore and refine interpretable concepts from clinical notes through expert feedback, thereby accelerating the development of more accurate and generalizable clinical prediction models while highlighting the critical role of human oversight in guiding the AI.

Original authors: Jean Feng, Avni Kothari, Patrick Vossler, Andrew Bishara, Lucas Zier, Newton Addo, Aaron Kornblith, Yan Shuo Tan, Chandan Singh

Published 2026-06-11
📖 5 min read🧠 Deep dive

Original authors: Jean Feng, Avni Kothari, Patrick Vossler, Andrew Bishara, Lucas Zier, Newton Addo, Aaron Kornblith, Yan Shuo Tan, Chandan Singh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build the ultimate, easy-to-use checklist for a doctor to predict if a patient is going to get sick. Traditionally, building this checklist is like trying to assemble a complex piece of furniture while blindfolded, with a team of experts arguing over which screws to use, how tight to turn them, and which instructions to follow. It takes forever, requires a massive team, and often, the final product is too complicated to actually use in a busy hospital.

This paper introduces a new way to build these checklists called HACHI. Think of HACHI as a "Human + AI" dance partnership. Instead of the humans doing all the heavy lifting and the AI just sitting there, they work together in a loop where the AI does the heavy exploring, and the humans steer the ship.

Here is how the HACHI partnership works, using simple analogies:

The Problem: The "Infinite Library"

Doctors have access to millions of pages of unstructured notes (like handwritten scribbles or typed paragraphs) about patients. These notes contain an "infinite library" of potential clues (concepts) about what makes a patient sick.

  • The Old Way: A human team tries to read through these notes and guess which clues matter. They might miss the most important ones or pick ones that look good on paper but don't work in real life.
  • The AI-Only Way: You could let a super-smart AI (a Large Language Model) read all the notes and pick the clues. But the AI might get confused, pick weird clues (like "the doctor wrote in blue ink" instead of "the patient has a fever"), or miss the big picture because it doesn't understand the "why" behind the medicine.

The Solution: The HACHI Dance

HACHI solves this by having the AI and the humans take turns in a loop.

  1. The AI Agent (The Scout): The AI is sent into the "infinite library" of notes. Its job is to rapidly scan, find potential clues, and turn them into simple Yes/No questions (e.g., "Does the note say the patient has a headache?"). It then tests these questions to see which ones help predict the sickness best.
  2. The Human Team (The Coaches): The humans look at the AI's results. They act like coaches. They might say:
    • "Hey, you picked a clue about 'note-writing style,' but that's cheating. Let's only look at the patient's actual symptoms."
    • "You missed the surgery type! That's a huge factor. Go look for that."
    • "This clue is too vague. Let's make the question more specific."
  3. The Loop: The humans give this feedback, and the AI goes back out to find better clues based on the new instructions. They repeat this a few times (usually 3 or 4 rounds) until the checklist is perfect.

Real-World Examples from the Paper

The team tested this on two specific medical problems:

1. Predicting Brain Injuries in Kids (TBI)

  • The Goal: Decide which kids with head injuries need a CT scan (which uses radiation) and which don't.
  • The AI's First Mistake: The AI initially thought "Does the note mention the Glasgow Coma Scale?" was a good clue. The humans said, "No! That's just a clue about how the doctor wrote the note, not the patient's condition." They also found the AI was accidentally using data from patients who had already been scanned at other hospitals (data leakage).
  • The Fix: The humans told the AI to ignore note-writing styles and remove those specific patients.
  • The Result: The final checklist was simple, accurate, and even found a new clue the old methods missed: "Does the child have a normal walking gait?" (If a child walks normally after a head hit, they are likely fine).

2. Predicting Kidney Injury After Surgery (AKI)

  • The Goal: Predict if a patient will get kidney damage after general surgery.
  • The AI's First Mistake: The AI only looked at the patient's body (like their weight or blood pressure) and ignored the surgery itself.
  • The Fix: The humans told the AI, "Stop! The type of surgery matters just as much. Also, make your questions more specific so different doctors interpret them the same way."
  • The Result: The final checklist included things like "Is the surgery urgent?" and "Is the patient having a heart rate over 100?" It became much more accurate than previous models.

Why This Matters

The paper claims that HACHI is faster and better than the old ways because:

  • It's Fast: The team only spent about 1–2 hours per round reviewing the AI's work.
  • It's Clear: The final models are just simple lists of Yes/No questions that any doctor can understand and use without a computer.
  • It's Trustworthy: Because humans are in the loop, they catch the AI's mistakes (like data leaks or weird correlations) before the model is finished.

In short, HACHI is like having a super-fast, super-smart intern (the AI) who does all the research, while the senior doctors (the humans) guide the intern to make sure the final report is accurate, fair, and actually useful. The paper shows that this partnership creates better medical prediction tools than either the humans or the AI could build alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →