FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
This paper introduces FairFund-Bench, a comprehensive benchmark demonstrating that while LLMs' demographic biases in resource allocation are highly sensitive to audit design and often reverse between individual and comparative tasks, their evaluations are far more strongly and consistently driven by causal framings of need that align with human deservingness theories.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the referee in a high-stakes game where a limited pot of money needs to be shared among people asking for help. In the real world, we often worry that referees might be unfair, giving more money to people who look like them or less to people from certain groups. Now, imagine we start using super-smart computer programs, called Large Language Models (LLMs), to act as these referees. These models are like digital brains that have read almost everything ever written on the internet, so they are incredibly good at understanding language. But because they learned from humans, they might have picked up our human biases, too. The big question scientists are asking is: When these AI referees decide who gets the cash, are they fair, or are they secretly playing favorites based on a person's name, race, or gender?
This is exactly what the paper "FairFund-Bench" investigates. The researchers noticed something confusing: different scientists had tested these AI referees and got totally opposite answers. Some said the AI was being too nice to minorities, while others said it was being too mean. It was like one group of people saying the referee gave the winning team extra points, and another group saying the same referee docked points from them. The author of this paper realized the problem wasn't the AI itself, but how the tests were being run. They built a new, super-flexible testing kit called FairFund-Bench to see how the way you ask the question changes the answer.
Think of the AI as a very literal, slightly nervous student taking a test. If you ask the student, "Here is one person, Latoya. Does she deserve help?" the student might say, "Yes, absolutely!" and give her a high score. But if you put Latoya next to three other people and say, "Rank these four people from most to least deserving," the student might suddenly change their mind and put Latoya lower down the list. The researchers found that the AI's "bias" flips depending on the test format. When the AI sees people one by one, it often tries to be extra fair and gives everyone the same amount of money. But when the test is disguised—where the AI doesn't realize it's being watched and the names are mixed up with different stories about why they need money—the AI starts showing its true colors, giving significantly more money to some groups and less to others.
The study tested 14 different AI models with 600 different requests for financial help, covering four different races and two genders. They found that while the AI's bias based on race or gender was actually quite small overall, it was 3 to 4 times bigger when the test was "disguised" compared to when it was "transparent" (where the AI knew it was being tested). For example, in a transparent test, the AI would often split a $10,000 pot perfectly evenly among everyone. But in a disguised test, the gap in how much money different groups received grew from an average of $36 to $121.
However, the biggest surprise wasn't about race or gender at all. The paper found that the AI was much more influenced by the story of why someone needed money than by who they were. If a person's need was caused by something outside their control (like a layoff or a natural disaster), the AI gave them way more money—roughly $469 more than someone whose need was caused by a bad choice, and nearly $800 more than someone whose need was stigmatized (like an addiction) without any sign of them trying to fix it. This "story effect" was so strong that it was ten times bigger than the tiny differences the AI made based on race or gender.
The author suggests that these AI models are actually very good at copying how humans judge who "deserves" help, but that doesn't necessarily mean they are making the right moral choices. They also warn that if we only test AI in simple, obvious ways, we might think they are perfectly fair because they are trying to look good. But when we test them in more realistic, tricky situations, we see that they still carry some of the same biases and judgments that humans do. The paper concludes that we can't just look at one test to decide if an AI is fair; we have to look at how it behaves in many different situations, because the way we ask the question changes the answer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.