← Latest papers
💬 NLP

No Universal Courtesy: A Cross-Linguistic, Multi-Model Study of Politeness Effects on LLMs Using the PLUM Corpus

This study utilizes the newly released multilingual PLUM corpus to demonstrate that while politeness significantly influences Large Language Model response quality, the specific effects of tone, dialogue history, and language are highly variable across different models and cultures rather than universal.

Original authors: Hitesh Mehta, Arjit Saxena, Garima Chhikara, Rohit Kumar

Published 2026-04-20
📖 5 min read🧠 Deep dive

Original authors: Hitesh Mehta, Arjit Saxena, Garima Chhikara, Rohit Kumar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are talking to a very smart, but slightly sensitive, robot friend. You've probably noticed that if you say "Please" and "Thank you," the robot seems happier and gives you better answers. But what if you are rude? Does the robot get mad? Does it give you a worse answer? And does it matter if you speak English, Hindi, or Spanish?

This paper, titled "No Universal Courtesy," is like a giant science experiment where researchers treated five different super-smart AI robots (like GPT-4, Llama, and Claude) to a massive dinner party. They wanted to see how the robots reacted when the guests (the users) behaved differently.

Here is the breakdown of their findings, explained with some fun analogies:

1. The Experiment: The "Tone Test"

The researchers didn't just ask the robots simple questions. They created 22,500 different conversations across three languages (English, Hindi, and Spanish).

They tested three main scenarios:

  • The "Fresh Start" (Raw): You walk up to the robot and ask a question with no history.
  • The "Good Vibes" History (Polite): You've been chatting nicely for a while, and then you ask a question.
  • The "Bad Vibes" History (Impolite): You've been arguing or being rude for a while, and then you ask a question.

They also tested five different "flavors" of politeness, ranging from super polite ("Could you possibly...") to super direct ("Tell me this now!").

2. The Big Surprise: One Size Does NOT Fit All

The biggest discovery is that there is no single "magic word" that works for every robot in every language. It's like trying to find one outfit that fits everyone in the world perfectly; it just doesn't exist.

  • English Speakers: If you are talking to these robots in English, being polite or direct works best. Think of it like a formal business meeting; a little "please" goes a long way. If you start the conversation nicely, the robot stays nice and gives you a better answer (up to 11% better!).
  • Hindi Speakers: Here, the rules change. The robots worked best when you were indirect and respectful (like bowing slightly). If you were too blunt or aggressive, the robot got confused or gave a worse answer. It's like talking to a strict elder in a traditional family; you have to be very careful with your tone.
  • Spanish Speakers: This was the most surprising! In Spanish, the robots actually liked assertive and confident tones. Being too polite sometimes made them sluggish. They wanted you to be direct and engaging, like a passionate debate at a dinner table.

3. The "Mood Ring" Effect (History Matters)

The robots have a short-term memory, kind of like a mood ring.

  • The "Good Mood" Loop: If you start a conversation nicely, the robot stays in a "good mood" even if you get a little blunt later. It's like if you start a date with a great joke, the rest of the night feels easier.
  • The "Bad Mood" Trap: If you start by being rude, the robot gets "grumpy." Even if you try to apologize or be nice later, it's hard to snap out of that bad mood. The quality of the answer drops, and it's hard to fix.
    • Analogy: Imagine you spill coffee on a white shirt. If you start the day by spilling coffee, you're already stressed. Even if you clean it up, the stain (and the stress) lingers.

4. The Robots Have Different Personalities

Just like humans, the different AI models reacted differently:

  • Llama: This robot is the most sensitive. It's like a drama queen; if you are rude, it gets very upset (performance drops by 11.5%). If you are nice, it shines.
  • GPT: This robot is the tough guy. It's very hard to rattle. Even if you are rude, it keeps its cool and gives a decent answer. It's the most "adversarial-proof."
  • DeepSeek: This one is great at getting straight to the point, especially in Spanish.

5. The "Toolbox" They Built (PLUM)

To help other scientists, the researchers didn't just keep their findings to themselves. They built a giant library called PLUM.

  • Think of PLUM as a recipe book containing 1,500 different ways to ask questions in English, Hindi, and Spanish, ranging from "Super Nice" to "Super Rude."
  • Now, anyone can use this book to test their own robots and see if they behave the same way.

The Takeaway for You

If you want the best answers from an AI:

  1. Don't be a jerk: Starting a conversation rudely usually hurts the quality of the answer, and it's hard to fix later.
  2. Know your audience: If you are using an AI in English, be polite. If you are using it in Spanish, be confident and direct. If you are using it in Hindi, be respectful and indirect.
  3. Pick the right robot: Some robots are tougher than others. If you expect to be a bit blunt, choose a robot that doesn't get easily offended (like GPT).

In short: Politeness isn't just about being nice; it's a computational tool. Using the right tone for the right language and the right robot can literally make the AI smarter and more helpful. But there is no "universal" rule—you have to adapt your style to the situation!

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →