GPT-4o and DeepSeek-V3 responses to common migraine questions: a cross-sectional comparative study of accuracy, completeness, empathy, readability and safety
This cross-sectional comparative study found that while both GPT-4o and DeepSeek-V3 provided generally accurate and safe information on migraine, DeepSeek-V3 demonstrated significantly higher scores in completeness and readability, though neither model was uniformly superior in empathy or safety.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Millions of people live with migraine, a condition that goes far beyond a simple headache to cause significant disability and disrupt daily life. Because the symptoms of a migraine can sometimes look like those of a stroke or other serious brain conditions, patients often turn to the internet for answers. In recent years, a new type of computer program known as a large language model has become a popular source for this information. These systems can read vast amounts of text and generate human-like answers to almost any question, offering immediate explanations about diagnosis, triggers, and treatments. However, because these programs are not doctors, there is a real concern that they might provide incomplete advice, miss dangerous warning signs, or offer reassurance when a person actually needs urgent medical care.
To understand how well these digital tools perform in a high-stakes medical setting, researchers conducted a direct comparison between two of the most advanced systems available: GPT-4o and DeepSeek-V3. The study focused specifically on the common questions people ask about migraine. The researchers gathered one hundred distinct questions from real-world sources, including online search trends and discussions on public forums where patients share their experiences. These questions covered a wide range of topics, from the nature of warning symptoms and lifestyle triggers to medication safety and long-term outcomes. To ensure the evaluation was fair and unbiased, two neurologists reviewed the answers generated by each model without knowing which computer program had produced them. They rated the responses based on how accurate the medical facts were, how complete the information was, whether the tone was empathetic and supportive, and most importantly, whether the advice was safe for a patient to follow.
The results showed that both computer programs generally provided accurate and safe information, successfully recognizing the need for professional medical help when serious warning signs were mentioned. However, one system, DeepSeek-V3, consistently outperformed the other in specific areas. It provided more complete answers, ensuring that patients received not just the basic facts but also necessary context about alternatives and uncertainties. Furthermore, the text generated by DeepSeek-V3 was significantly easier to read and understand, using simpler sentence structures and clearer language. While both models showed similar levels of empathy in their tone, and both were safe enough to be considered useful tools, the difference in completeness and readability was clear and statistically significant.
In a specific test involving twenty-five questions that described dangerous "red flag" symptoms, such as a sudden worst-ever headache or new neurological problems, both models performed well by correctly identifying the need for urgent care. Yet, even in these critical scenarios, the system that produced the clearer, more readable text maintained its advantage. The researchers concluded that while these artificial intelligence tools are becoming powerful allies for patient education, they should never replace a doctor's assessment. Instead, they are best used to help people understand general concepts and recognize when they need to seek professional help, provided their answers are reviewed with care. The study highlights that as these technologies evolve, their ability to communicate complex medical ideas simply and completely is just as important as the accuracy of the facts they present.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.