← Latest papers
💬 NLP

OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants

This paper introduces OATH-Frames, a nine-category typology for analyzing online attitudes toward homelessness in the U.S., and demonstrates that leveraging large language models to assist domain experts can scale the annotation of 2.4 million Twitter posts with a 6.5x speedup and minimal performance loss, yielding nuanced insights into public sentiment across different demographics and regions.

Original authors: Jaspreet Ranjit, Brihi Joshi, Rebecca Dorn, Laura Petry, Olga Koumoundouros, Jayne Bottarini, Peichen Liu, Eric Rice, Swabha Swayamdipta

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Jaspreet Ranjit, Brihi Joshi, Rebecca Dorn, Laura Petry, Olga Koumoundouros, Jayne Bottarini, Peichen Liu, Eric Rice, Swabha Swayamdipta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to understand how millions of people feel about a complex, emotional issue like homelessness. If you tried to read every single tweet about it, you'd be buried under a mountain of text faster than you could blink. It's like trying to drink from a firehose.

This paper, titled "OATH-Frames," is about building a better pair of glasses to see what's actually in that firehose. The researchers used a mix of human social work experts and powerful AI (Large Language Models) to sort through 2.4 million tweets about homelessness in the U.S.

Here is the story of how they did it, broken down into simple parts:

1. The Problem: The "One-Size-Fits-All" Glasses

Before this study, if you wanted to analyze tweets about homelessness, you might use standard tools that just ask: "Is this tweet mean?" (Toxicity) or "Is this tweet happy or sad?" (Sentiment).

The researchers argue these tools are like trying to sort a mixed bag of fruit using only a "Red" and "Green" sticker.

  • The Flaw: A tweet might say, "Homeless people are lazy and dirty." Standard tools might miss this because it's not using "swear words" (low toxicity) and it's not "happy" (negative sentiment). But it is deeply harmful.
  • The Reality: People's attitudes are messy. You can feel sympathy for a homeless person and be angry at the government for not fixing it, all in the same sentence. Simple "good/bad" labels miss this nuance.

2. The Solution: The "OATH-Frames" Menu

The team created a new "menu" of nine specific categories (frames) to describe exactly how people are talking about homelessness. Think of this like a chef's tasting menu instead of just "food."

They grouped these nine categories into three main flavors:

  • Critiques (The Complaints):
    • Government Critique: Blaming politicians or laws.
    • Societal Critique: Blaming society or "hypocritical" people.
    • Money Aid: Arguing about where tax dollars should go.
  • Perceptions (The Stereotypes):
    • Harmful Generalizations: Calling all homeless people thieves, addicts, or lazy.
    • Deserving vs. Undeserving: Arguing that some groups (like veterans) deserve help more than others (like immigrants).
    • Not In My Backyard (NIMBY): "I don't mind homeless people, just not in my neighborhood."
  • Responses (The Actions):
    • Solutions: Suggesting fixes, shelters, or policy changes.
    • Personal Interaction: Sharing a story about meeting a homeless person.
    • Media Portrayal: Talking about how the news or TV shows homelessness.

3. The Process: The "Human-AI Dance"

Sorting 2.4 million tweets by hand would take a human team years. Sorting them with AI alone is risky because AI can misunderstand subtle social cues (like sarcasm or specific political references).

So, they invented a "Human-AI Dance":

  1. The Experts: Social work experts first defined the rules (the "menu") and labeled a small batch of tweets to teach the AI what to look for.
  2. The AI Assistant (GPT-4): The AI tried to label thousands of tweets. When it got confused, it explained its reasoning (like a student showing their work).
  3. The Check-Up: The human experts looked at the AI's "work" and fixed the mistakes. They found that by letting the AI do the heavy lifting and just checking the results, they got 6.5 times faster at labeling, with only a tiny drop in accuracy.
  4. The Final Stretch: Once the AI learned the ropes from the experts, they trained a smaller, faster AI model to label the remaining 2.4 million tweets on its own.

4. What They Found: The Hidden Patterns

Once they had all 2.4 million tweets sorted into these nine categories, they could see patterns that were previously invisible.

  • The "Toxicity" Blind Spot: They found that many tweets containing harmful stereotypes (like calling homeless people "dirty") had low "toxicity" scores. The AI's new "OATH-Frames" caught these, while the old tools missed them.
  • Location Matters:
    • In California, Washington, and Oregon, people were more likely to use Harmful Generalizations (stereotypes). The researchers think this is because there are more visible, unsheltered homeless populations there, leading to more direct (and sometimes negative) reactions.
    • In New York, people were more likely to argue about Deserving vs. Undeserving groups. This seemed linked to news about migrants and asylum seekers arriving in the city, sparking debates about who "deserves" aid.
  • The "Ukraine vs. Immigrant" Split:
    • When people compared homeless Americans to Ukrainians, the tweets were mostly about Government Critique and Money Aid (e.g., "Why are we sending billions to Ukraine when we have homeless people here?").
    • When people compared homeless Americans to Immigrants, the tweets were much more hostile, filled with Harmful Generalizations and NIMBY attitudes (e.g., "They are taking our jobs/housing").
  • Time Matters: During big news events (like the Russian-Ukraine war funding debates), the tweets spiked in specific categories. When Congress was discussing aid for Ukraine, the "Money Aid" and "Deserving vs. Undeserving" arguments exploded in the tweets.

The Bottom Line

This paper didn't just count tweets; it taught a computer how to understand the nuance of human anger, sympathy, and confusion regarding homelessness.

By combining human social work expertise with AI speed, they created a tool that can spot subtle, harmful attitudes that standard "mean vs. nice" detectors miss. They showed that how we talk about homelessness changes depending on where we live, who we are comparing them to, and what's happening in the news right now.

Note: The paper focuses entirely on analyzing these online attitudes to help researchers and advocacy groups understand public opinion better. It does not claim to solve homelessness directly or offer clinical advice, but rather provides a clearer lens to see the problem.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →