Clinically Grounded AI-Scribing in Psychotherapy: Benchmarking LLMs Against Expert Documentation in the iCARE Framework
This paper introduces the clinically grounded iCARE documentation framework and the corresponding iHOPE dataset to benchmark eleven large language models, revealing that while closed-source models generally outperform open-source ones in automated metrics, human expert evaluation highlights specific strengths and weaknesses—such as struggles with temporal reasoning and varying clinical preferences—that underscore the need for AI tools designed to assist rather than replace therapists.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the quiet rooms where psychotherapy takes place, the work is deeply human. It relies on a therapist listening to a client's story, noticing the weight of a pause, and understanding the context of a fear that stretches back years. For the therapist to do this work effectively, they must also keep a record. These notes are not merely administrative paperwork; they are the map of a person's inner world, capturing everything from the immediate crisis to the long history of family and habits that shaped them. Traditionally, writing these notes has been a heavy burden, pulling the clinician's attention away from the patient and into a computer screen. In recent years, artificial intelligence has promised to lift this burden. These "AI scribes" can listen to a conversation and type it out, but in the complex world of mental health, most current systems fall short. They tend to produce brief, dry summaries that miss the nuance, the emotional undercurrents, and the critical safety details that define a therapy session. The challenge has been to build a machine that does not just transcribe words, but understands the story.
A team of researchers from India and Singapore has taken a significant step toward solving this problem by rethinking how therapy notes should be structured and how machines should learn to write them. They began by creating a new, comprehensive framework for documentation called iCARE. Unlike the standard, abbreviated formats used in many hospitals, which are designed for quick administrative checks, iCARE is built to capture the full depth of a clinical encounter. It is a detailed structure with seventeen distinct sections, ranging from basic identifying information and the client's chief concerns to a deep dive into their medical history, a risk assessment for immediate dangers like self-harm, and a plan for future sessions. The researchers argued that an AI system must first be capable of generating this rich, complete picture before it can be condensed into a shorter summary. To test this idea, they built a new dataset called iHOPE. They took 174 recorded therapy sessions from a public collection, cleaned up the audio to ensure every word was heard, and then had a team of expert psychiatrists and psychologists write perfect, gold-standard notes for each session using the new iCARE format. This created a benchmark, a set of ideal answers against which any machine could be measured.
The researchers then put eleven different large language models to the test. These models, which are the powerful computer programs behind modern AI chatbots, were asked to listen to the therapy recordings and write the iCARE notes. They tested the models in two ways: first, by giving them the recordings with no examples to follow, and second, by showing them one example of a perfect note before asking them to write the rest. The results revealed a clear divide between the different types of AI. The closed-source models, which are developed by major technology companies and are not open for public modification, consistently performed better than the open-source models, which are built by the public community. One specific model, GPT-4o-mini, achieved the highest scores on standard computer tests that measure how closely the machine's words matched the human experts' words. Another model, Gemini Pro, showed a particular strength in identifying crisis markers, such as signs of suicide risk or substance abuse, which are critical safety details in therapy.
However, the story became more interesting when the researchers asked human experts to judge the notes, rather than relying on computer scores. They brought in six mental health professionals to read the notes blindly, without knowing which AI wrote them, and rate them on five key qualities: trustworthiness, relevance, accuracy, completeness, and clarity. The human experts found that while the top computer models were good, they still struggled with the flow of time. The machines often failed to correctly distinguish between what happened in a past session and what was planned for the next one, a fundamental aspect of therapy that requires understanding a timeline. Surprisingly, the human experts preferred the notes from a smaller, open-source model called Mistral over the more powerful commercial models. In fact, Mistral was the only system that scored slightly higher than the human-written notes in one specific category: completeness. This finding suggests that the way we currently measure AI success with computer algorithms does not always match what a human clinician actually needs. The commercial models were excellent at matching specific phrases, but the smaller model seemed to capture the clinical essence of the session more effectively in the eyes of the experts.
The study concludes that while artificial intelligence is ready to assist therapists, it is not yet ready to replace them. The researchers found that the best path forward is a collaborative one, where AI handles the heavy lifting of drafting the detailed notes, allowing the therapist to review, refine, and finalize the record. This approach could free up valuable time for clinicians to focus on the patient rather than the keyboard. The team emphasized that their work is not a final solution but a foundation. They have made their new framework, their dataset of expert notes, and their evaluation method available to other scientists so that the field can continue to improve. The ultimate goal is to bridge the gap between the massive need for mental health care and the limited number of available therapists, using technology to support, rather than substitute, the human connection at the heart of healing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.