← Latest papers
📄 medicine

An agentic AI workflow for spine history taking and surgeon-facing clinical synthesis

In a prospective single-clinic study, the SpineGPT-2 agentic AI system demonstrated that safety-first, automated history-taking can significantly reduce physician consultation time while producing clinically adequate, surgeon-facing summaries with high diagnostic alignment and red-flag detection accuracy.

Original authors: Begüm Aslantaş Kaplan, Ali Kaplan, Ali Aydilek, Özkan Çeliker, Tolga Ege, Salim Şentürk, İlker Solmaz

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Begüm Aslantaş Kaplan, Ali Kaplan, Ali Aydilek, Özkan Çeliker, Tolga Ege, Salim Şentürk, İlker Solmaz

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quiet, often crowded waiting rooms of spine clinics, a fundamental challenge plays out every day: the need to gather a complete medical story before a surgeon can make a safe decision. For patients with back or neck pain, the history they tell is not just a formality; it is the most powerful tool doctors have to distinguish between simple muscle strain and serious conditions like infections, tumors, or nerve damage. Yet, the reality of modern medicine often means doctors are pressed for time, forced to rush through these critical conversations. When a history is incomplete, the path to a correct diagnosis can become unclear, leading to unnecessary tests or delayed care. This tension between the need for deep, structured questioning and the pressure of a busy schedule has led researchers to ask a new question: can a computer help? Specifically, can an artificial intelligence system, designed to listen and ask the right follow-up questions, gather this complex medical history before the patient ever sees the doctor, and then summarize it in a way that a surgeon can trust?

To answer this, a team of researchers in Turkey developed a system called SpineGPT-2. Unlike a simple chatbot that just answers questions, this system was built as an "agentic" workflow, a term that describes a computer program capable of acting with a specific goal in mind. The system operates in three distinct stages. First, a patient-facing agent conducts a text-based interview with the patient, asking one question at a time and adapting its next question based on the answer given, much like a skilled interviewer who follows a thread of conversation. Second, a rigid, rule-based computer layer sits in the background, ensuring the interview stays safe and covers all necessary medical topics without letting the computer make dangerous mistakes or give medical advice. Finally, a surgeon-facing agent takes the full transcript of that conversation and synthesizes it into a structured report, highlighting key symptoms, potential diagnoses, and safety warnings for the doctor to read before the patient enters the room.

The researchers tested this system in a real-world setting at a single spine clinic. They invited 240 patients to participate, splitting them into two groups based on alternating days. On some days, patients used the SpineGPT-2 system to complete their history interview while waiting. On other days, patients underwent the standard process where the surgeon asked all the questions directly. The goal was to see if the computer-generated summary was good enough to be useful. An independent panel of expert surgeons, who were aware that the summaries they were reviewing were system-generated, rated them on a scale to determine if the information was sufficient to make confident clinical decisions.

The results showed that the system worked remarkably well. The expert panel rated the average quality of the computer-generated summaries as 4.23 out of 5, a score that significantly exceeded the pre-set threshold of 4.0 required for the summaries to be considered safe and actionable. In nearly 77 percent of the cases, the panel agreed that the summary was good enough to proceed with the consultation without needing to repeat the entire history. The system also proved to be highly accurate in identifying the correct medical territory. In 88 percent of cases, the panel's leading diagnosis appeared within the top three possibilities suggested by the computer. Furthermore, the system was very good at spotting "red flags"—symptoms that suggest a serious underlying condition like a fracture or infection. It correctly identified these warning signs in 93.8 percent of the cases where they were present, and it rarely raised false alarms when they were absent.

Beyond accuracy, the system offered a tangible benefit to the flow of the clinic. When the surgeon had a pre-written summary to review, the time they spent asking the patient for their history dropped significantly. On average, the doctor's history-taking time fell from just over ten minutes in the standard group to less than seven minutes in the group using the system. This saved time did not come at the cost of patient satisfaction; those who interacted with the computer agent found the experience acceptable and usable. However, the researchers were careful to note that the system is not perfect. In about 8 percent of cases, the summary contained small errors, such as a symptom mentioned that the patient never actually reported, or a minor mistake in how a piece of information was linked to a medical guideline. Crucially, none of these errors involved missing a critical safety warning, but the presence of these small mistakes reinforced the need for a human doctor to always review the final report.

The study concludes that this type of agentic AI system can successfully act as a supervised assistant in a busy spine clinic. It does not replace the surgeon's judgment or the physical examination, but it handles the repetitive, time-consuming task of gathering the initial story, organizing it, and presenting it clearly. By shifting the burden of the initial interview to a reliable digital agent, the system allows the human doctor to focus their time on verification, physical examination, and the complex decision-making that requires human experience. The findings suggest that with the right safety guardrails and human oversight, artificial intelligence can move from being a theoretical tool to a practical partner in improving how spine care is delivered.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →