"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
This paper introduces a novel multi-domain corpus and audit protocol for tracking social biases against homelessness across online and offline discourse, revealing that while large language models can identify such biases, they suffer from significant miscalibration that systematically over-flags "NIMBY" sentiments and under-detects factual claims.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, chaotic town square where millions of people shout their thoughts, complaints, and stories every second. Now, imagine that some of these shouts are about a very serious, painful problem: people who have nowhere to sleep. This is the world of "homelessness," a situation where people lack a safe place to live. In this digital town square, people argue about where to build shelters, whether homeless people deserve help, or why they are even there. Sometimes, these arguments are helpful; other times, they are filled with prejudice, fear, or mean-spirited stereotypes.
Scientists who study language on computers (a field called Natural Language Processing) have built "smart robots" called Large Language Models (LLMs). Think of these robots as super-fast librarians who can read millions of books, tweets, and forum posts in a blink. Their job is often to act as a "bias detector," scanning the noise to find specific types of mean or unfair speech. But here's the catch: just because a robot is fast and has read a lot doesn't mean it understands the feeling or the context of what it's reading. It might mistake a question for an attack, or think that mentioning a house means someone hates homeless people. This paper asks a crucial question: Can these digital librarians actually tell the difference between a genuine complaint and a hateful bias, or are they just guessing based on the wrong clues?
The "Not in My Backyard" Mix-Up
In this study, a team of researchers from the University of Notre Dame and the United Nations University decided to test these smart robots on a very specific, sticky topic: the bias against people experiencing homelessness (PEH). They wanted to see if the robots could accurately spot when people were saying things like "Not In My Backyard" (NIMBY)—a fancy way of saying, "I don't want a homeless shelter built near my house."
To do this, they built a massive digital library. They gathered over 50,447 pieces of text from ten different U.S. cities, pulling from four very different places: Reddit threads, X (formerly Twitter) posts, news articles, and actual transcripts from city council meetings where real people argue about policy. They covered a whole decade, from 2015 to 2025.
But a library is only as good as its catalog. So, the researchers didn't just let the robots label everything. They brought in human experts—people who work with homeless populations in non-profit organizations—to read a carefully selected sample of 1,698 posts. These humans acted as the "gold standard," the truth-tellers who decided what each post actually meant. They used a detailed checklist with 16 different categories to label the text, ranging from "expressing an opinion" to "harmful generalizations" and, of course, "Not In My Backyard."
The Robot's Big Mistake
The researchers then asked six different smart robots (including big names like GPT-4.1, Gemini, and LLaMA) to read the same posts and guess the labels. They expected the robots to be pretty good, maybe getting about 43% of the tricky labels right (a score called "macro-F1"). And sure enough, the robots looked okay on paper. They seemed to be doing a decent job.
But then, the researchers looked closer, and that's when they found the glitch.
They discovered that the robots were suffering from a massive "calibration error." Imagine a smoke detector that is so sensitive it screams "FIRE!" every time you toast a piece of bread. That's what happened here. Every single robot the team tested was way too eager to label a post as "Not In My Backyard."
On average, the robots over-claimed NIMBY bias by 11.5 percentage points compared to the human experts. In plain English: if a human said, "This person is just stating a fact," the robot often shouted, "This person is a NIMBY!"
For example, if someone posted, "They've got all that affordable housing, but the downside is it's in [City]," the humans labeled this as just "stating a fact" or "expressing an opinion." But four out of the six robots immediately tagged it as "Not In My Backyard." Why? Because the robots got tricked by two simple clues:
- Housing words: If the text mentioned "housing" or "affordable housing," the robot assumed someone was against it.
- Question marks: If the post was a question, the robot assumed it was a sarcastic, rhetorical attack.
The robots were essentially playing a game of "guess the bias" based on a shallow dictionary of words, rather than actually understanding the meaning. They treated the mention of housing as proof of opposition to housing.
The "Fact" Blind Spot
The robots didn't just over-claim bias; they also missed the truth. The study found that the robots were terrible at spotting when someone was actually providing a fact or a claim. They under-detected factual statements by a whopping 30.5 percentage points.
This is a dangerous mix. If you use these robots to monitor public opinion for a city council, you would end up with a distorted picture of reality. You would think the community is much more hostile and opposed to shelters than they actually are (because of the fake NIMBY alarms), and you would think they are offering much less evidence-based information than they really are (because the robots ignored the facts).
What This Means for the Future
The researchers didn't just point out the problem; they offered a way to fix it. They realized that you can't just trust the robots' raw numbers. Even if a robot says it's "80% accurate," it might still be lying about how many people are biased.
The paper suggests that if cities or organizations want to use these AI tools to track stigma against homeless people, they must run a "prevalence audit." This means constantly checking the robot's work against a small, trusted group of human experts to see if the robot is inflating the numbers. They also found that smaller, local robots (like the ones you could run on your own computer for privacy) were even worse at this than the big, expensive ones.
In the end, this paper is a warning label for the digital age. It shows that while AI is a powerful tool for listening to the world, it is currently a very clumsy listener. It hears the words "housing" and "question mark" and jumps to the conclusion that someone is being mean. Until we teach these robots to understand the nuance between a fact and a fear, we can't let them run the show on their own. We need to keep the human experts in the loop to make sure we aren't just listening to the robot's imagination.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.