← Latest papers
💻 computer science

Evaluating Clinical Symptom Extraction and Reasoning in Local Large Language Models: A Proof-of-Concept Study of Symptom-Specific Performance in Cardiology

This proof-of-concept study evaluates locally deployed large language models (Gemma3 and Qwen3.5) on Japanese cardiology discharge summaries, revealing that while overall symptom extraction accuracy is high, performance varies by specific symptom and the models often infer conditions like palpitations or fatigability from objective clinical findings rather than explicit patient complaints.

Original authors: Shun Kitamura, Eisuke Amiya, Yoshihiro Izawa, Risa Kishikawa, Takenobu Shimada, Junichi Ishida, Satoshi Kodera, Norihiko Takeda

Published 2026-09-17
📖 6 min read🧠 Deep dive

Original authors: Shun Kitamura, Eisuke Amiya, Yoshihiro Izawa, Risa Kishikawa, Takenobu Shimada, Junichi Ishida, Satoshi Kodera, Norihiko Takeda

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Hospitals generate vast amounts of written records every day, from the initial description of a patient's pain to the final notes on their recovery. Much of this information is locked inside free-flowing paragraphs of text, making it difficult to sort, search, or analyze quickly. For decades, doctors and researchers have relied on manual effort to read these notes and pull out specific details, a slow and tedious process that is prone to human error. Recently, a new type of computer program known as a large language model has emerged as a potential solution. These systems are trained on enormous libraries of text and can understand human language well enough to read a medical note and identify specific facts, such as whether a patient mentioned feeling short of breath or experiencing chest pain. Because these tools can work entirely within a hospital's own secure computers without sending private data to the outside world, they offer a promising way to turn messy handwritten or typed notes into clean, organized data that can be used to improve patient care.

A team of researchers at The University of Tokyo set out to test how well two of these local computer programs could perform this task in the specific and high-stakes field of heart failure. They focused on patients with advanced heart failure who were being evaluated for a heart transplant, a group whose symptoms are complex and varied. The researchers chose ten real patient records from their own hospital, covering the period between 2017 and 2023. These records were written in Japanese and contained the full story of each patient's hospital stay, including what they complained about, what the doctors found during exams, and how their condition changed over time. The team asked the computer programs to look at these ten stories and identify the presence or absence of eight specific symptoms required for transplant listing, such as palpitations, extreme tiredness, and fainting. To ensure the computer's answers were accurate, five expert heart specialists independently reviewed the same ten stories and agreed on a "ground truth" for each symptom, creating a reliable standard against which to measure the machines.

The researchers tested the computer programs under different sets of instructions to see how the wording of the request changed the results. In one scenario, the programs were told to scan the entire medical history for any mention of the symptoms. In another, they were explicitly told to ignore any symptoms that might have been mentioned in the patient's past history and to focus only on what was happening during this specific hospital admission. They also tested whether asking the programs to "think out loud" and explain their reasoning before giving an answer would improve their accuracy. The study found that both computer programs were generally quite good at the task, correctly identifying the presence or absence of symptoms in most cases. However, the performance was not perfect, and the errors were not random. The programs struggled most with two specific symptoms: palpitations, which is the sensation of a racing or pounding heart, and fatigability, or an overwhelming sense of tiredness.

When the researchers examined the mistakes the computers made, they discovered a clear pattern in how the machines were thinking. For the symptom of palpitations, the programs often decided a patient was experiencing them simply because the medical record showed an abnormal heart rhythm or a fast heart rate, even if the patient never actually wrote or said they felt their heart racing. The computer inferred the symptom from the objective medical data rather than from a direct patient complaint. Similarly, for fatigability, the programs frequently concluded that a patient was tired because the patient had heart failure, a condition known to cause tiredness, even when the specific note did not mention fatigue at all. The computer was filling in the gaps with medical logic rather than sticking strictly to the text. This tendency was so strong that in some cases, the computer correctly identified that a patient had an irregular heartbeat but still falsely claimed the patient felt palpitations, while in other cases with the same medical findings, it correctly said the patient did not feel them, showing an inconsistency in its reasoning.

The study also revealed that the way the instructions were phrased mattered significantly. When the researchers told one of the programs to ignore past medical history and focus only on the current hospital stay, the program became much more careful about finding symptoms, but it also missed many that were actually present. It became so strict about ignoring anything that wasn't explicitly written in the current context that it failed to catch symptoms that were clearly part of the patient's current condition. This suggests that while these computer programs are powerful tools, they do not simply read words like a human; they interpret the text based on medical associations and patterns they have learned. They can infer that a symptom exists based on other facts, which can be helpful but also leads to errors when that inference is wrong. The researchers noted that these findings are based on a small number of cases and that the programs behaved differently depending on the specific instructions given, indicating that the technology is still evolving.

Ultimately, this work shows that while local computer programs can automate the extraction of medical information, they do not yet replace the careful judgment of a human reader. The machines are capable of understanding the context of a medical note, but they sometimes let their knowledge of how diseases work override what is actually written on the page. For these tools to be truly useful in a hospital setting, developers will need to understand exactly how the programs reason through a diagnosis and how to guide them to rely on explicit patient complaints rather than medical assumptions. The study serves as a proof of concept, demonstrating that these systems can work within a hospital's secure network and handle complex medical text, but it also highlights the need for careful oversight to ensure that the automated extraction of symptoms remains accurate and reliable.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →