← Latest papers
💻 computer science

Auditing Sex/Gender Disparities in Emergency Triage with LLM-based Paired Comparisons

This paper introduces an LLM-based paired-comparison method to audit sex/gender disparities in emergency triage documentation, revealing that identical clinical presentations are slightly more likely to receive lower severity scores when labeled as female, a finding that validates the approach as a scalable tool for generating bias hypotheses rather than confirming direct clinical undertriage.

Original authors: Ariel Guerra-Adames, Marta Avalos-Fernandez, Océane Dorémus, Leo Anthony Celi, Cédric Gil-Jardiné, Emmanuel Lagarde

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Ariel Guerra-Adames, Marta Avalos-Fernandez, Océane Dorémus, Leo Anthony Celi, Cédric Gil-Jardiné, Emmanuel Lagarde

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but instead of looking for fingerprints, you are looking for invisible biases hidden inside the way people make decisions. This paper lives in the world of Artificial Intelligence (AI) and Healthcare, specifically focusing on how we can use smart computer programs to check if doctors and nurses are treating patients fairly. The core idea relies on a concept called a "counterfactual," which is just a fancy way of asking, "What would have happened if things were slightly different?" In this case, the question is: "If this patient were a different gender, but had the exact same symptoms, would they get a different level of care?" We care about this because in emergency rooms, seconds count. If a patient is sent to the back of the line because of who they are rather than how sick they are, it could be dangerous. Scientists have long suspected that gender might play a hidden role in these split-second decisions, but it's very hard to prove because we can't rewind time and ask a nurse, "What if this was a man?"

This study uses a new kind of AI tool, called a Large Language Model (LLM), to act as a "bias mirror." Think of the LLM as a super-smart robot that has read millions of medical records and learned to mimic how human nurses decide how urgent a patient is. The researchers didn't build this robot to replace doctors; they built it to be a doctor for a moment, so they could test it. They took over 140,000 real emergency room stories from a hospital in France and a database in the US. Then, they used the AI to create "twin" versions of these stories. For every story about a woman, the AI wrote a new version where the patient was a man, but the symptoms, the pain, and the history were kept exactly the same—like swapping a red shirt for a blue one while keeping the person inside identical. They did the same for men turning into women.

The results were like finding a tiny, consistent crack in a perfectly smooth mirror. The study found that when the AI looked at these "gender-swapped" twins, it tended to give the female version a slightly lower urgency score than the male version. In plain English, the robot thought, "This person is less sick because they are a woman," even though the symptoms were identical. The numbers were small but clear: in the French data, a female presentation was about 1.1% more likely to get a less severe score than the male version. In the US data, that gap was about 2.2%. It's like if you and your twin brother both had a broken arm, and the robot triage nurse decided your brother needed to see a doctor sooner just because he's a boy.

The researchers were careful to check if this was just a glitch in the robot's brain or if it was actually reflecting something in the real data. They tried two things: first, they checked if the robot was just "hallucinating" (making things up), and second, they checked if the bias came from the robot's general training or the specific medical data. They found that when they scrubbed the data of all gender words and numbers before teaching the robot, the bias disappeared. This suggests the robot wasn't making up the difference; it was faithfully copying a pattern it saw in the real medical records. Interestingly, the bias seemed strongest when the nurse and the patient were the same gender (like a female nurse treating a female patient), suggesting these patterns are complex and human, not just random errors.

However, the authors are very careful not to sound the alarm too loudly. They emphasize that these are small, documentation-level effects. The robot is looking at written notes, not the actual patient standing in front of a nurse. The study proves that the written records contain these gendered patterns, but it doesn't prove that every single patient was actually treated unfairly at the bedside. The authors call this a "feasibility study"—a proof that we can use AI to find these hidden signals. They suggest that if these small differences in the notes do translate to real life, it could mean thousands of women a year in France might get slightly delayed care. But until human experts re-check these specific cases, we can't say for sure how much harm is actually being done. The main takeaway is that we now have a powerful, scalable magnifying glass to find these invisible biases, and we know they are there, waiting to be fixed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →