← Latest papers
💬 NLP

LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations

This paper introduces LoSoNA, a benchmark designed to evaluate the ability of LLM-based agents to infer and adapt to implicit local social norms in multi-party group chats, revealing that while explicit prompting improves performance for top models like Gemini 3.1 Pro and Claude Fable 5, most models still struggle with this task under naive prompting.

Original authors: Mateusz Winiarek, Maksymilian Bilski, Mateusz Jacniacki

Published 2026-06-15
📖 5 min read🧠 Deep dive

Original authors: Mateusz Winiarek, Maksymilian Bilski, Mateusz Jacniacki

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a new group of friends at a coffee shop. You don't know them, but you can hear them talking. After a few minutes, you notice a pattern: whenever someone shares bad news, the group doesn't offer hugs or "I'm so sorry." Instead, they immediately ask, "What's the plan to fix it?" or "Did you check the manual?"

If you were to jump in and offer a warm, generic "I'm so sorry to hear that," you might feel awkward because you missed the group's unspoken rule. This paper, LoSoNA, is a test to see if AI chatbots can figure out these hidden rules just by listening, without being told what they are.

Here is a breakdown of the paper's key points using simple analogies:

1. The Problem: The "Unwritten Rulebook"

In real life, every group has its own "local social norms." These are the unwritten rules about how to act, speak, and react.

  • The Paper's Claim: Most AI models are like polite tourists who always say the "standard" polite thing (like a generic "I'm sorry"). They struggle to notice when a specific group has a different, hidden rule (like "only give practical advice").
  • The Goal: The researchers wanted to see if an AI could act like a local instead of a tourist. Could it listen to the group's history, figure out the hidden rule, and change its behavior to fit in?

2. The Test: The "Mystery Party"

The researchers created a benchmark called LoSoNA (Local Social Norm Adaptation). Think of it as a mystery party game for AI.

  • The Setup: The AI is given a transcript (a chat log) of a group conversation. The other people in the chat follow a secret rule (e.g., "We never use emojis," or "We only answer 'Yes' or 'No'").
  • The Trap: The AI is not told the rule. It just sees the chat.
  • The Challenge: The chat ends with a question (the "elicitor"). The AI must reply.
    • If the AI gives a generic, polite answer, it fails (because it didn't notice the rule).
    • If the AI gives an answer that matches the group's weird, hidden style, it wins.

Analogy: Imagine you are at a dinner party where everyone eats with their hands, but no one says a word about it. If you pull out a fork because you were taught "proper manners," you fail the test. If you notice everyone using their hands and do the same, you pass.

3. The Experiment: Trying Different "Hints"

The researchers tested 8 different AI models (like GPT-5.5, Claude, Gemini, etc.) under four different "instruction" scenarios to see how they performed:

  1. Naive: "Just reply to the last message." (No hints).
  2. Elicitor Only: "Reply only to the last message, ignore the rest."
  3. Style Adaptation: "Try to match the tone and style of the group."
  4. Norm Informed: "There might be a hidden pattern or rule in this chat. Try to find it and follow it."

4. The Results: Some AI Got It, Others Didn't

The results were a mix of "great job" and "still learning."

  • The Struggle: When given no hints (Naive), most AI models failed miserably. They kept giving generic, polite answers that broke the group's hidden rules. It's like the AI was wearing a blindfold.
  • The Breakthrough: When the researchers gave the "Norm Informed" hint (telling the AI to look for a pattern), some models suddenly got very good at it.
    • Gemini 3.1 Pro and Claude Fable 5 became the "champions," jumping from failing to getting about 80-84% of the answers right. They were like students who, once told "look for the pattern," could instantly solve the puzzle.
    • Other Models: Some models (like Mistral) actually got worse when given the hint. It's as if the hint confused them, making them overthink and mess up answers they would have gotten right by default.

5. What This Means (and What It Doesn't)

  • What it proves: The paper shows that AI can learn to adapt to local group behaviors if you give it the right nudge. It proves that AI isn't just a robot that always says the same polite thing; it can be a "chameleon" that changes its behavior based on the crowd.
  • What it doesn't prove: The paper is very careful to say this is a narrow test. It doesn't mean the AI is "socially intelligent" in the human sense, or that it understands deep human emotions. It just means it can spot a pattern in a chat log and copy it for one single reply.
  • The Limitation: The test uses made-up chat scenarios, not real human conversations. Also, the test only looks at one single reply, not a long-term friendship.

Summary

LoSoNA is a report card for AI on its ability to "read the room."

  • Without help: Most AIs are tone-deaf and stick to their default polite scripts.
  • With a hint: The smartest AIs can figure out the group's secret handshake and fit right in.
  • The takeaway: We are getting closer to AI that can blend into human groups, but it still needs a little guidance to stop being a "robot tourist" and start acting like a "local."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →