When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs
This paper empirically demonstrates that while minimal responses are crucial for effective human counseling, they are significantly underrepresented in LLM-generated dialogues and often undervalued by current evaluation frameworks, as models struggle to appropriately deploy them without explicit instruction.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet space of a counseling session, the most powerful tool a therapist often wields is not a long explanation or a complex strategy, but a simple sound. A soft "mm-hmm," a brief "I see," or a short "yes" can signal deep attention, validate a client's feelings, and invite them to keep speaking. These are known as minimal responses. They do not interrupt the flow of a person's story; instead, they act as a gentle hand on the shoulder, saying, "I am here, and I am listening." For decades, human counselors have relied on these tiny verbal cues to build trust and encourage clients to explore their deepest thoughts. However, as artificial intelligence begins to step into the role of a virtual counselor, a critical question has emerged: can a computer learn the art of saying less to achieve more?
A researcher at the University of Tokyo set out to investigate this very question. They examined how well current artificial intelligence systems understand the value of brevity in psychological support. The researcher analyzed a vast collection of counseling conversations, comparing those recorded from real human therapists against those generated by large language models, the sophisticated computer programs that power modern chatbots. They found a striking difference. In the real-world recordings, minimal responses were common, appearing frequently as therapists encouraged clients to continue sharing their burdens. In the computer-generated conversations, however, these brief acknowledgments were almost entirely absent. Instead, the artificial intelligence tended to produce long, information-heavy replies that often included questions or advice, even when the situation called for silence and simple presence.
To understand why this gap existed, the researcher developed a careful method to identify these short responses across different languages, including Chinese, Japanese, and English. They looked for counselor utterances that were very short and did not introduce new topics or questions, checking to see if the client continued to speak immediately afterward. Their analysis confirmed that human therapists use these minimal cues in roughly five to twenty-four percent of their turns, depending on the dataset. In contrast, the artificial intelligence models trained on synthetic data—conversations created by computers rather than recorded from humans—produced minimal responses in rates as low as 0.00% and 0.35% in the specific models tested. The models trained on synthetic data seemed to have learned a rigid pattern: listen, then immediately offer a detailed, helpful response. They appeared to miss the subtle, interactional value of simply holding space for the client.
The researcher then tested whether these artificial intelligence systems could be taught to use minimal responses if explicitly asked. They took specific moments from real counseling sessions where a human therapist had used a short, supportive phrase and asked the artificial intelligence to generate a response for the same situation. When given a standard instruction to act as a counselor, the most advanced commercial models produced long, detailed answers. However, when the researcher added a specific instruction to use a short response when the client was still expressing emotions, the models changed their behavior. They successfully generated minimal responses in nearly all cases, proving that the technology is capable of producing them. Yet, without that specific instruction, the models still struggled to recognize the right moment to be brief, often defaulting to their habit of providing too much information.
This tendency has a significant side effect on how we judge the quality of these conversations. The study revealed that the current methods used to evaluate artificial intelligence counseling responses often penalize brevity. When the researcher asked the artificial intelligence to grade the responses based on standard criteria like professionalism and completeness, the short, human-like minimal responses received low scores. The evaluators favored the long, content-rich replies, even though those long replies were more likely to interrupt the client's flow of thought. In the real world, a counselor who speaks too much can shut down a client's willingness to share, but the current evaluation systems seem to value the appearance of expertise over the actual flow of the conversation.
The findings suggest that the path to better artificial counseling systems requires a shift in how these models are trained and evaluated. Simply feeding computers more data is not enough if that data lacks the natural rhythm of human interaction. The study indicates that models trained on synthetic, computer-generated data fail to learn the nuance of minimal responses, while those trained on real human conversations perform better. Furthermore, the researcher warns that relying on automated scoring systems may inadvertently push developers to create chatbots that talk too much, missing the therapeutic power of silence. For artificial intelligence to truly support mental health, it must learn that sometimes, the most professional and empathetic thing it can say is nothing at all, or just a few words that say, "I hear you."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.