An Investigation of Linguistic Biases in LLM-Based Recommendations
This study investigates linguistic biases in LLM-based restaurant and product recommendations across Southern American, Indian, and Code-Switched Hindi-English dialects, revealing that specific model families exhibit heightened sensitivity to Indian English and Code-Switched prompts, leading to distinct variations in recommendation patterns without consistent trends based solely on model size.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of very smart, digital chefs (the AI models) who are experts at suggesting what you should eat or buy. You ask them the exact same question: "I'm hungry, what's good to eat?" or "I need a new shirt, what do you suggest?"
But here's the twist: you ask the question in three different "accents" or styles of speaking:
- Southern American English: Like a friendly neighbor in the South.
- Indian English: A specific way English is spoken in India.
- Code-Switched: A mix of English and Hindi, like "Arre yaar, I'm dying of hunger!"
The researchers wanted to see if these digital chefs would give you different answers just because of how you asked, even if what you asked was the same. They treated the AI like a "cold start" chef—meaning they didn't give the AI any history about your past likes or dislikes, just the question and a list of options.
The Experiment: The "Menu" Test
The researchers gave the AI a long, balanced list of 100 restaurants (half American, half Indian) and 100 products (like clothes, home goods, or beauty items). They asked the AI to pick its top 20 favorites based on the specific "accent" of the question. They did this hundreds of times with different lists to make sure the results were real and not just a fluke.
What They Found: The "Accent" Matters
The study found that the AI's "taste" changed depending on the dialect, much like a human might guess your preferences based on your accent.
1. The Restaurant Results:
- The Pattern: When asked in Indian English or Code-Switched (Hindi-English), the AI was much more likely to recommend Indian restaurants. When asked in Southern American English, it recommended fewer Indian restaurants.
- The "Big Brain" Effect: The largest AI model tested (Llama-3.1-70B) was the most sensitive to these accents. It was like a chef who, upon hearing a specific accent, immediately thought, "Ah, this person must want Indian food," and served up a plate of it, even though the question didn't explicitly say "I want Indian food."
- The "Smartest" Chef: Interestingly, the GPT-OSS models (which are known for being very good at reasoning) were the least affected by the accent. They seemed to ignore the "flavor" of the words and focus more on the actual meaning, acting like a chef who says, "You said you're hungry; here is a balanced meal from the whole menu," regardless of how you asked.
2. The Product Results:
- The Pattern: The AI also changed its shopping suggestions based on the accent.
- Code-Switched prompts often led to more recommendations for Home, Sports, and Clothing items.
- Indian English prompts often led to more Beauty product recommendations.
- Size Doesn't Always Fix It: Usually, we think bigger AI models are "smarter" and less biased. But here, the bigger models didn't always behave better. Sometimes the big models were more sensitive to the accent than the small ones, and sometimes they were less. There was no simple rule that "bigger is better" for fixing this bias.
The Big Takeaway
The paper concludes that these AI systems are not just reading the meaning of your words; they are also reacting to the style of your words.
Think of it like a waiter who, upon hearing a specific accent, assumes you want a specific type of cuisine and brings it out before you even order. While this might sometimes be helpful, the researchers warn that it can be a problem. It means the AI is making assumptions about who you are and what you like based on your language, which can limit your choices and reinforce stereotypes.
In short: If you ask an AI a question in a different dialect, it might give you a different list of recommendations, even if you asked for the exact same thing. The AI is listening to your "voice," not just your "words."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.