Automatic detection of depression in clinical interviews with large language models
This study demonstrates that automated language analysis of clinical interviews using transformer-based models and ensemble strategies can effectively detect major depressive disorder in Cantonese speakers, significantly improving processing efficiency while highlighting the importance of transcription quality and conversational context for optimal performance.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Depression is a common and serious condition that affects millions of people worldwide, yet finding it early often depends on a person's ability to access a trained mental health professional. In many places, these experts are scarce, and the process of diagnosis can be slow, leaving many without the help they need. To address this, researchers have begun exploring whether computers can learn to spot signs of depression by listening to how people speak. This field relies on the idea that the way a person describes their feelings, the words they choose, and the rhythm of their conversation can reveal their mental state. Recent advances in artificial intelligence have created powerful tools capable of understanding human language, offering a new way to analyze these conversations automatically. The goal is not to replace doctors, but to create a fast, accessible tool that can help identify who needs care, potentially reaching people who would otherwise go undiagnosed.
A team of researchers at the Chinese University of Hong Kong set out to test whether these artificial intelligence tools could accurately detect major depressive disorder using recordings of clinical interviews. They worked with 299 Cantonese-speaking participants, including 194 patients diagnosed with depression and 105 healthy individuals. The researchers recorded structured interviews where a clinician asked questions and the participant responded. To see if they could make the process faster and more scalable, they compared two ways of turning these audio recordings into text. The first method was the traditional approach, where humans manually wrote down every word spoken, a process that took about 45 minutes for each interview. The second method used an automated system to transcribe the audio instantly. They then fed these texts into several different computer models designed to understand language, testing whether the models could tell the difference between the healthy participants and those with depression.
The study found that using the automated transcription system was a massive leap forward in efficiency. It reduced the time needed to process each interview from roughly 45 minutes down to just under 10 minutes, a 79 percent reduction. This speed makes it possible to analyze large numbers of interviews quickly, which is essential for real-world use. However, the researchers discovered that this speed came with a small trade-off. The models trained on the manually written texts were slightly better at identifying depression than those trained on the automatically generated texts. This suggests that while machines are getting very good at listening, they still occasionally miss the subtle nuances that a human transcriber catches. Despite this small gap, the automated versions were still highly effective, proving that the technology is ready for practical application where speed is critical.
Another key discovery was about what parts of the conversation the computer should listen to. The researchers tested two scenarios: one where the model only read the patient's answers, and another where it read the entire conversation, including the doctor's questions. They found that the models performed better when they had the full context of the dialogue. A patient's short answer, such as "not really" or "sometimes," might seem vague on its own, but when the model sees the question that came before it, the meaning becomes much clearer. For instance, if a doctor asks about sleep patterns and the patient says "sometimes," the computer understands this refers to sleep disturbances. This context helps the model make more accurate judgments, showing that the interaction between the doctor and the patient holds valuable clues that are lost if only the patient's voice is analyzed.
To make their detection even more reliable, the team combined the predictions of several different computer models into a single, stronger system. They tested various ways of merging these opinions, such as taking the average of their scores or looking for the strongest signal among them. This combined approach, known as an ensemble, improved the accuracy of the detection beyond what any single model could achieve on its own. The best-performing combination was then tested on a completely new group of 169 young people who had not been part of the original study. In this independent test, the combined system successfully identified depression with a high level of accuracy, performing just as well as the best single model and slightly better than the others. This confirmed that the system could generalize its findings to new people, not just the ones it was trained on.
The results of this study offer a promising path forward for mental health care. The researchers demonstrated that it is possible to build a system that detects depression from speech with high accuracy while drastically cutting down the time required for analysis. Although the automated transcripts were not perfect, they were good enough to support a system that works well in the real world. The study also highlighted that including the doctor's questions in the analysis provides crucial context that improves the computer's understanding. While the system is not yet a replacement for a human diagnosis, it represents a significant step toward making mental health screening faster, more accessible, and available to more people who need it. The work suggests that with the right tools, technology can help bridge the gap between those suffering from depression and the care they deserve.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.