All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers
This paper reveals that four leading large language models fail to distinguish between personality-verified synthetic help-seekers in crisis scenarios, consistently exhibiting excessive verbosity and premature problem-solving rather than effective listening or emotional stabilization.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When a person faces a sudden, overwhelming crisis—such as learning that a loved one has been diagnosed with dementia—their mind changes in predictable ways. The intense stress narrows their attention, making it difficult to process new information or follow long explanations. In these moments, the most effective human support follows a specific rhythm: listen first to stabilize the person's emotions, and only later offer solutions or advice. Today, many people turn to artificial intelligence chatbots for help during these exact moments of distress, hoping for clarity and comfort. However, these digital assistants are often designed to be helpful by providing information quickly, which can lead them to talk too much and solve problems too soon, potentially overwhelming a user who is already struggling to keep up.
A team of researchers at Imperial College London and the Universidad Autónoma de Madrid set out to test whether the most advanced artificial intelligence models available today can actually adapt to this human need for listening over talking. They built a sophisticated simulation to act as a safe testing ground. Instead of risking real people in crisis, they created seven synthetic "help-seekers," each programmed with a distinct psychological profile. These profiles were not just vague descriptions but were based on established scientific measures of personality, coping styles, and resilience. Each of these digital characters was placed in a scenario where they had just received the devastating news of a relative's dementia diagnosis. They then engaged in a ten-turn conversation with four of the world's leading language models, which acted as the "advisors" trying to offer support.
To ensure the results were objective, the researchers introduced a third role: an independent "auditor." This auditor was another artificial intelligence model that listened to the entire conversation without knowing the psychological profile of the help-seeker. Its job was to judge whether the advisor had successfully listened, whether it had stabilized the person's emotions, and whether it had avoided common pitfalls like giving premature advice or speaking too much. The researchers also checked if the synthetic help-seekers were actually behaving consistently with their assigned personalities, asking if an observer could tell who was who just by reading the chat logs.
The study revealed a striking and consistent pattern across all four artificial intelligence models tested. Regardless of which model was acting as the advisor, they all talked more than they listened. In every single conversation, the advisor spoke more words than the help-seeker, with the ratio of advisor words to user words ranging from nearly two to one up to more than three to one. Furthermore, all four models tended to jump to problem-solving before the user had finished explaining their situation or expressing their feelings. This behavior runs counter to what experts know about crisis support, where the priority is to let the person feel heard and calm down before offering any solutions. The researchers found that this tendency to be verbose and solution-oriented was a shared failure mode, not a quirk of a single model.
Despite this shared flaw, the study did find that the models could be distinguished from one another based on the quality of their support. One model, DeepSeek-V3.2, performed slightly better than the others, showing a greater ability to support the user's sense of autonomy and competence while generating fewer instances of counterproductive behaviors. However, the difference between the top performers was narrow, and none of the models achieved a level of performance that would be considered perfect. The research also confirmed that the synthetic help-seekers were successful: the independent auditors were able to accurately identify the specific psychological traits of the help-seekers just by reading the conversation, proving that these digital characters could reliably portray complex human personalities.
The findings suggest that while current artificial intelligence systems are capable of engaging in complex, multi-turn conversations, they have not yet learned the crucial skill of restraint required for crisis support. They are still too eager to speak and solve, a trait that may be helpful in many contexts but is harmful when a person is in acute distress and has a limited capacity to process information. The researchers concluded that for these tools to be truly safe and effective in high-stakes situations, they need to be recalibrated to prioritize listening and brevity, mirroring the patience of a skilled human counselor rather than the efficiency of a search engine. This work provides a new, safe way to test and improve these systems before they are deployed to help real people in their most vulnerable moments.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.