← Latest papers
📄 health informatics

Quality of physicians responses using a conversational artificial intelligence system: a randomized vignette experiment in Latin America

In a randomized vignette experiment involving physicians in Latin America, an evidence-traceable conversational AI system significantly improved the composite validity of clinical responses compared to usual search practices, though findings are considered exploratory due to the volunteer-based sample.

Original authors: Castano-Villegas, N., Monsalve, K., Villa, M. C., Quiros Gomez, O. I., Velasquez, L., Zea, J.

Published 2026-09-25
📖 4 min read☕ Coffee break read

Original authors: Castano-Villegas, N., Monsalve, K., Villa, M. C., Quiros Gomez, O. I., Velasquez, L., Zea, J.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the bustling clinics of Latin America, a doctor often faces a moment of quiet uncertainty: a patient presents with a complex set of symptoms, and the correct treatment path is not immediately obvious. For decades, the solution has been to pause and search for answers, flipping through textbooks or typing queries into a digital database. This process relies on the physician's ability to find, read, and interpret vast amounts of medical information quickly. Today, a new tool has entered this landscape: conversational artificial intelligence. These systems, built on massive collections of human knowledge, can hold a dialogue, answer questions, and summarize complex topics in seconds. The promise is that such a tool could act as a partner at the bedside, helping doctors retrieve evidence faster and more accurately. However, most tests of these systems have been conducted in controlled, static environments, far removed from the reality of a busy hospital in a low- or middle-income country. The critical question remains whether these digital assistants actually improve the quality of care when used by real doctors in real settings, or if they are merely sophisticated chatbots that sound confident but offer little practical help.

To find the answer, researchers in Latin America designed a direct comparison involving two hundred and two physicians. They created a scenario where doctors were asked to solve four simulated patient cases, each requiring answers to four open-ended questions. The doctors were randomly split into two groups. One group used their usual methods to find information, searching through standard resources as they normally would. The other group used a specific conversational system designed to provide answers that could be traced back to medical evidence. To ensure fairness, two independent specialists, who did not know which method each doctor had used, graded every response. They evaluated the answers based on six different dimensions of quality, creating a composite score that reflected the overall validity of the medical advice given.

The results showed a clear difference between the two approaches. The physicians who used the conversational system produced answers that were, on average, more valid than those who relied on their usual search practices. The scores for the supported group were higher, with a median of 2.83 compared to 2.46 for the usual practice group. When the researchers adjusted for factors like the doctors' academic degrees, the data suggested that a physician using the system was significantly more likely to produce an answer that met a high standard of validity. This advantage held true across most of the specific criteria used for grading, particularly those where the expert judges agreed most strongly with one another. However, the study also revealed a nuance: when looking strictly at the accuracy of the facts alone, the difference between the two groups did not reach a level of statistical significance. This led the researchers to designate the broader, composite measure of validity as the primary finding, a decision reinforced when an external evaluator repeated the accuracy-only analysis and found the same lack of difference.

Beyond the quality of the answers, the study examined how the doctors felt about the process and how long it took them. The physicians who used the system reported high levels of satisfaction, finding the tool acceptable and easy to use. Interestingly, the time it took to complete the tasks and the number of searches the doctors reported performing did not show a clear difference between the two groups, though the researchers noted that missing data made this specific comparison difficult to interpret with certainty. The study concludes with an important reminder about its scope: the participants were volunteers who chose to complete the exercise, meaning the findings are exploratory rather than definitive proof of how the tool works for every doctor in every situation. While the system did not act as a magic wand that solved every problem instantly, the evidence suggests that when used by practicing physicians in this region, it can help elevate the quality of the medical information they provide to their patients.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →