Do Language Models Know Their Slang? Queer Slang Understanding in User-Generated Content
This paper introduces Slang-Q, a manually curated dataset and taxonomy of 118 queer slang terms, to evaluate the ability of language models to accurately understand and define this underrepresented community-specific language.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling city where every neighborhood has its own secret handshake, inside jokes, and special vocabulary. In the world of science, there's a field called Natural Language Processing (NLP) that builds digital brains—known as Large Language Models (LLMs)—to read and understand this city. Think of these models as super-smart tourists who have read almost every book and website ever written, but they sometimes struggle to understand the local slang of specific neighborhoods. One such neighborhood is the LGBTQIA+ community, which has a rich, creative, and rapidly changing language full of words that can mean very different things depending on who is saying them and where they are. If these digital tourists get the slang wrong, they might misunderstand a friendly greeting as an insult, or miss the point of a joke entirely. This matters because these models are increasingly used to help people find information, chat with friends, and navigate the web; if they don't understand the language of marginalized communities, they can't serve everyone fairly.
Enter a team of researchers who decided to test just how well these digital tourists know the "Queer Slang" neighborhood. They built a special evaluation ground called Slang-Q, which is like a massive, carefully organized dictionary of 118 specific queer terms paired with 1,024 real-life sentences where people actually used them. Crucially, this list of words was created independently of the sentences, ensuring the models hadn't just memorized the answers. They didn't just ask the models to guess; they set up a series of challenges to see if the models could define these words correctly. The researchers tested four different ways of asking the models: sometimes they just gave the word (like asking "What is 'tea'?"), and sometimes they gave the word plus a sentence showing how it was used (like "I'm spilling the tea about the party"). They also tried two different "masks" for the models: one where the model acted like a generic language expert, and another where the model was told to act specifically as an expert in queer internet slang.
The results were a bit like watching a student take a test without studying the right textbook. When the models were asked to define the slang terms without any context or special instructions, they often stumbled. They tended to pick the most common, everyday meaning of a word instead of the specific queer meaning. For example, if you asked about "bear," the model might think of a furry animal instead of a specific type of person in the gay community. However, the story changed when the researchers gave the models a little help. When they provided a sentence showing how the word was used, or when they told the model, "Hey, you are a queer slang expert," the models got much better at guessing the right meaning. It's as if the models were like a person who knows the dictionary definition of a word but needs a little context clue to remember how their friends actually use it at a party.
The researchers found that while the models are getting smarter, they still aren't perfect. Even with the best hints, the models didn't quite reach the level of a human expert who grew up in the community. The models often missed the subtle cultural nuances, like whether a word was originally an insult that the community reclaimed to mean something positive, or if it had a specific connection to certain cultural groups. Interestingly, the study also checked if the models had just memorized the answers from their training data. They found that some models, particularly the commercial ones, seemed to have "seen" these specific sentences before, while the open-source ones hadn't. But even the ones that hadn't memorized the answers could figure them out if given the right context.
In the end, this paper suggests that while our digital brains are powerful, they still need a little guidance to truly understand the vibrant, shifting language of the queer community. They can't just rely on a generic dictionary; they need to be told, "This is a special neighborhood, and here is how the locals talk." The study doesn't claim to have solved the problem forever, but it does show that giving models the right context and framing makes a huge difference. It's a reminder that to build truly helpful AI, we need to teach it not just the words, but the culture and the stories behind them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.