← Latest papers
🤖 AI

Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents

This paper introduces a seven-dimensional taxonomy to empirically evaluate 49 studies on LLM-based agents in healthcare, revealing significant asymmetries where information-centric capabilities are well-developed while critical action-oriented tasks, safety mechanisms, and adaptive learning features remain largely underimplemented.

Original authors: Shubham Vatsal, Harsh Dubey, Aditi Singh

Published 2026-02-05
📖 4 min read☕ Coffee break read

Original authors: Shubham Vatsal, Harsh Dubey, Aditi Singh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a team of 49 different "digital doctors" (AI agents powered by Large Language Models) that researchers have built to help with healthcare tasks. Some are great at reading patient charts, others are good at answering medical questions, and a few are trying to figure out how to plan treatments.

The authors of this paper decided to stop just listing what these robots can do and instead built a 7-point checklist (a taxonomy) to grade them fairly. They wanted to see which skills are common, which are missing, and where the whole field stands today.

Here is the breakdown of their findings, explained simply:

The 7-Point Checklist

The researchers looked at every "digital doctor" through seven different lenses:

  1. Brain Power (Cognitive Capabilities): Can the AI plan ahead, understand complex inputs, and fix its own mistakes?
  2. Library Access (Knowledge Management): Does it just rely on what it memorized during training, or does it know how to look up fresh information (like new drug guidelines) in real-time?
  3. Conversation Style (Interaction Patterns): Does it just wait for a command, or does it listen for alarms (like a sudden drop in heart rate) and know when to ask a human for help?
  4. Learning Ability (Adaptation & Learning): If the world changes (e.g., a new virus appears), can the AI learn and update itself, or does it stay stuck in the past?
  5. Safety & Ethics: Does it have guardrails to stop it from giving dangerous advice? Is it fair to all patients, and does it keep secrets?
  6. Team Structure (Framework Typology): Is it a "lone wolf" doing everything alone, or a team of specialists working together?
  7. Job Description (Core Tasks): What specific medical jobs is it actually doing?

The Big Reveal: The "One-Sided" Robot

The most interesting finding is that these AI agents are very lopsided. They are like a student who is a genius at reading a textbook but terrible at taking a test or handling a real-life emergency.

  • The Superpower (What they are good at):

    • Looking things up: Most of them (about 76%) are excellent at grabbing external information (like medical guidelines) to answer questions.
    • Reading Charts: They are getting pretty good at summarizing patient records and answering medical trivia.
    • Teamwork: Most are built as teams of specialists (e.g., one agent reads the X-ray, another checks the drug list) rather than a single brain doing everything.
  • The Weaknesses (What they are missing):

    • The "Alarm Clock" Problem: Almost none of them (92% are missing this) can automatically wake up when a patient's vitals change. They mostly just sit there waiting for a human to type a question.
    • The "Forgetting" Problem: While they are good at looking up new info, they are terrible at forgetting old, outdated info. Only 2% of them have a system to actively delete old data to make room for new, correct data.
    • The "Safety Net" Problem: Very few have a built-in "panic button" to catch their own errors or a system to ask a human doctor to double-check their work before acting.
    • The "Adaptation" Problem: If the rules of medicine change, these agents generally don't notice. They are static; they don't learn from their mistakes in real-time.

The Job Market: What Can They Actually Do?

The researchers mapped out exactly what jobs these agents are currently hired for:

  • The "Librarians" (Common): They are great at Medical Question Answering and Document Summarization. If you need to know what a drug does or summarize a 100-page chart, these agents are ready.
  • The "Apprentices" (Partial): They are okay at Triage (deciding how urgent a problem is) and Diagnostic Reasoning, but they often need a human to hold their hand.
  • The "Strangers" (Rare): They are very bad at Treatment Planning (deciding exactly what medicine to prescribe) and Drug Discovery. These tasks require too much risk and precision, and the agents aren't there yet.

The Bottom Line

Think of current AI in healthcare as a brilliant intern who has read every medical book ever written but has never seen a real patient.

They are fantastic at retrieving facts and organizing information. However, they are not yet ready to be the "Chief of Staff" who makes the final call on a treatment plan, manages a crisis, or adapts to a changing situation without a human supervisor watching over them. The paper concludes that before we can trust these agents in real hospitals, we need to build better safety nets, better learning systems, and better ways for them to react automatically to emergencies.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →