← Latest papers
📄 medicine

Evaluating LLMs for HIV Epidemiological Surveillance in Spanish: Clinical Summarization, Risk Factor Extraction, and Inference Stability across Models

This study demonstrates that Large Language Models, particularly Gemini, are effective and stable tools for automating Spanish-language HIV epidemiological surveillance tasks, including clinical summarization and risk factor extraction, thereby offering a viable solution to reduce documentation burdens and improve early diagnosis in high-volume emergency settings.

Original authors: Ignacio Jolín-Rodrigo, Ana Carrero-Fernández, Juan Cuadros-González, Maria D. R-Moreno

Published 2026-09-04
📖 5 min read🧠 Deep dive

Original authors: Ignacio Jolín-Rodrigo, Ana Carrero-Fernández, Juan Cuadros-González, Maria D. R-Moreno

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the bustling, high-pressure environment of a hospital emergency room, doctors are constantly racing against time. They must piece together a patient's story from fragmented notes, scattered test results, and hurried observations to make life-saving decisions. For patients with HIV, a virus that weakens the body's immune system, this process is especially critical. Identifying risk factors early—such as unprotected sex, shared needles, or specific symptoms—can lead to testing and treatment that stops the virus from spreading. However, when a doctor is overwhelmed by dozens of patients a day, crucial details can easily slip through the cracks. The sheer volume of unstructured text in medical records makes it difficult to manually extract every relevant clue, creating a bottleneck where missed information can delay diagnosis and allow the disease to continue its silent spread.

To address this, researchers are exploring the use of advanced computer programs known as large language models. These are artificial intelligence systems trained on vast amounts of text that can read, understand, and summarize human language. The idea is that these tools could act as a second pair of eyes, rapidly scanning thousands of medical records to highlight key details and summarize patient histories. But for this to work in a real hospital, the technology must be incredibly reliable. It cannot invent facts, leave out vital information, or behave unpredictably. Furthermore, because these tools are often being tested in Spanish-speaking countries, they must understand the specific medical language and cultural nuances of Spanish clinical notes, not just English.

A team of researchers in Spain recently put this idea to the test. They wanted to see if these artificial intelligence programs could accurately summarize emergency room records and pull out specific HIV risk factors from Spanish text without needing to be retrained on new data first. They gathered 100 anonymized patient records from a university hospital, a mix of people who were later found to have HIV and those who were not. The records varied greatly in length and complexity, reflecting the messy reality of daily emergency care. The researchers then fed these records into eleven different AI models, ranging from powerful, cloud-based systems to open-source programs that could be run on a hospital's own secure servers. They asked the models to do two things: write a short, clear summary of the patient's visit, and identify whether specific risk factors were present, absent, or simply not mentioned in the text.

The results revealed a clear winner in the task of summarization. One model, a cloud-based system from Google, produced summaries that were significantly more accurate and complete than an open-source alternative. When doctors reviewed these summaries, they found that the cloud-based model made far fewer mistakes. It rarely invented facts that weren't there, and it missed far fewer important details. The open-source model, while capable, frequently left out critical information, such as pending test results or past medical history, which could be dangerous if a doctor relied on the summary alone. The researchers also tested whether an AI could judge the quality of these summaries as well as a human doctor. They found that the AI judge was very good at spotting made-up facts and obvious errors, but it struggled to agree with human doctors on what constituted a missing detail. This suggests that while AI can help filter out bad summaries, human oversight remains essential for catching subtle omissions.

When it came to extracting specific risk factors, the performance of the models depended heavily on how strictly the results were measured. The researchers ran each model five times on the same records to see if it would give the same answer every time. They found that the models were generally better at identifying social risk factors, like drug use or sexual history, which are often stated directly by patients. However, they were less consistent when identifying clinical risk factors, such as specific symptoms or past infections, which require piecing together scattered clues from different parts of a long medical note. Even when the models were set to be as consistent as possible, they sometimes gave different answers on repeated runs. This instability meant that a model that looked excellent in a single test might perform differently in a real-world, continuous workflow.

The study concluded that while these AI tools show great promise for helping Spanish-speaking emergency rooms manage their workload, they are not yet perfect. The best models can significantly reduce the time doctors spend writing summaries and can help surface risk factors that might otherwise be missed. However, the researchers emphasized that the technology must be used with caution. The most common error was not making things up, but rather leaving out important information. For a system designed to catch missed diagnoses, leaving out a detail is just as dangerous as inventing one. The findings suggest that the most effective approach is a hybrid one: using AI to handle the heavy lifting of reading and summarizing, while keeping human doctors in the loop to verify the results and ensure no critical piece of the puzzle is lost. This balance could help emergency departments catch more HIV cases earlier, potentially saving lives and slowing the spread of the virus in communities where it remains a persistent challenge.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →