← Latest papers
💻 computer science

Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment

This study reveals that while Large Language Models exhibit higher intellectual humility than human experts, they still reproduce human decision inconsistencies and a reductionist bias in neurodevelopmental disorder assessments by prioritizing biological survival over social needs, highlighting the need to critically evaluate the conceptual frameworks AI operationalizes in high-stakes mental health decisions.

Original authors: Maciej Wodziński, Joanna Wodzińska, Kacper Dudzic, Marcin Moskalewicz

Published 2026-08-19
📖 6 min read🧠 Deep dive

Original authors: Maciej Wodziński, Joanna Wodzińska, Kacper Dudzic, Marcin Moskalewicz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the complex world of mental health care, doctors and psychologists often face a difficult task: deciding if a person's condition is severe enough to qualify for government support. This is not just about diagnosing an illness; it is about judging how much a person can manage their own life. In many places, this judgment determines whether a family receives money for care, specialized housing, or the right to have a parent stay home to help. The decision hinges on a specific question: can this person live independently, or do they need constant help with their basic needs? For people with neurodevelopmental conditions like autism or attention deficit hyperactivity disorder, the answer is rarely simple. Their struggles might be invisible in a doctor's office but overwhelming in the real world. Because these decisions carry such weight, there is a growing hope that artificial intelligence could help make them fairer and more consistent. Large language models, the powerful computer programs that can write and reason, are being tested to see if they can act as impartial judges. But before we trust them with such high-stakes choices, we must understand how they see the world. Do they interpret human needs the same way humans do, or do they miss the subtle, social parts of a life that make independence possible?

A team of researchers in Poland set out to test this by putting both human experts and artificial intelligence through a rigorous series of trials. They gathered thirty-five human professionals, a mix of eighteen physicians and seventeen psychologists, who regularly work on disability assessment boards. These are the committees that decide who gets support. To compare them fairly, the researchers also selected seven different large language models from leading technology companies. The goal was to see if the machines would make the same mistakes as the humans, or if they would bring a new kind of bias to the table. The researchers designed a study that looked at three specific things: whether the decision-makers could be tricked by the order of information or by a misleading photo, how humble they were about their own knowledge, and, most importantly, how they defined the phrase "basic life needs."

The experiment began with a series of fictional case studies. The human experts and the computer models were asked to read descriptions of people with various conditions and then answer two questions. First, they had to rate how well the person was functioning. Second, they had to decide if that person needed constant care or assistance to live independently. To test if the decision-makers were easily swayed by irrelevant details, the researchers played a few tricks. In some cases, they placed all the negative information about a person's struggles at the very beginning of the story, hoping the reader would get stuck on those bad first impressions. In other cases, they attached a photo of a person looking sad or distressed to a story that was otherwise identical to one with a neutral photo. They wanted to see if seeing a sad face would make the experts grant more support, even if the text said the person was doing fine.

The results of these tricks were surprisingly consistent across both groups. Neither the human experts nor the artificial intelligence models showed a significant change in their decisions based on the order of the text or the emotion in the photo. Both groups seemed to look past these distractions and focus on the core facts of the case. This was a relief, suggesting that the models were not easily fooled by the same visual or textual cues that might trip up a human. However, when the researchers looked deeper into the actual decisions, a different kind of problem emerged. They found that neither the humans nor the machines were very consistent in their logic. A person could be rated as having a "significantly impaired" level of functioning, yet that same person might not be granted the support they needed. Conversely, someone rated as having "average" functioning might still get the help. The rating of how well a person was doing did not reliably predict whether they would get the support. This inconsistency existed in both the human doctors and the computer models, suggesting that the leap from describing a problem to granting a solution is a messy, human-like process that the machines have also inherited.

Perhaps the most revealing part of the study came when the researchers asked the participants to explain their thinking. They asked the humans and the machines to define what "basic life needs" actually meant. Here, a clear divide appeared. The human psychologists tended to have a broad view. They included emotional well-being, the ability to communicate, and social connection as essential parts of a basic life. If a person could not talk to others or understand social cues, the psychologists saw that as a failure to meet a basic need. The human physicians, however, took a narrower view. They focused almost entirely on physical survival: the ability to eat, dress, walk, and use the bathroom. The artificial intelligence models, surprisingly, sided with the physicians. They defined basic needs almost exclusively in terms of biological survival and self-care. They rarely mentioned communication or social interaction unless it was directly tied to physical safety.

This difference in definition had a direct impact on the final decisions. Because the psychologists included social and communicative struggles in their definition of "basic needs," they were much more likely to grant support to the fictional patients. The physicians and the artificial intelligence models, sticking to a strict biological checklist, granted support less often. Even though the models claimed to be very humble about their knowledge—scoring higher on a test of intellectual humility than the human doctors—this humility did not make their decisions more consistent or more aligned with the human experts. In fact, the models' high scores on humility seemed to be a surface-level trait, a way of speaking that did not translate into a deeper understanding of human complexity. The machines were willing to say they might be wrong, but they still operated on a narrow, medical framework that missed the social reality of the people they were judging.

The study concludes that while artificial intelligence might be good at ignoring simple tricks like a sad photo or a confusing sentence order, it is not yet ready to replace human judgment in these sensitive areas. The models are reproducing a reductionist view of human life, one that values physical independence over social and communicative connection. By focusing only on the body's ability to function, the machines risk denying support to people whose greatest struggles are in the realm of interaction and understanding. The researchers suggest that for artificial intelligence to be truly useful in mental health and disability assessment, it must be taught to see the whole person, not just the biological machine. Until the technology can grasp the full weight of what it means to live a human life, relying on it for these decisions may simply reinforce the same narrow perspectives that have long complicated the work of human experts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →