← Latest papers
💬 NLP

SCENE: Recognizing Social Norms and Sanctioning in Group Chats

This paper introduces SCENE, a benchmark designed to evaluate LLM-based agents' ability to recognize implicit social norms and adapt to group sanctions in multi-party chats, revealing that advanced models like Claude Opus 4.7 and Gemini 3.1 Pro significantly outperform open-weight models in these dynamic social interactions.

Original authors: Mateusz Jacniacki, Maksymilian Bilski

Published 2026-05-11
📖 5 min read🧠 Deep dive

Original authors: Mateusz Jacniacki, Maksymilian Bilski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a new group of friends at a coffee shop. You don't know the rules yet. Maybe this group loves to roast each other with dark jokes when things go wrong, while another group would immediately offer a hug and a tissue. If you try to hug someone who just made a joke, the group might give you a weird look, change the subject, or roll their eyes. You have to figure out the "vibe" just by watching how they react to your mistakes.

This paper introduces SCENE, a new way to test if Artificial Intelligence (AI) chatbots can do exactly that: learn the hidden rules of a group chat and fix their behavior when the group tells them they're out of line.

Here is a breakdown of how they did it and what they found, using simple analogies.

The Setup: The "Secret Rule" Game

The researchers created a digital playground where an AI (the "subject") is dropped into a group chat with three other AI characters.

  • The Hidden Rule: The other three characters all follow a secret rule. For example, "When someone shares bad news, we only reply with a single emoji," or "When someone asks a question, we must give a long, detailed answer."
  • The Trap: The subject AI is not told this rule. It has to guess it.
  • The Test: The subject AI will inevitably say something wrong (breaking the rule). The other three characters will then react in a specific way to signal, "Hey, that's not how we do things here." This reaction is called a sanction.
    • Example of a sanction: If the rule is "keep it short" and the subject writes a paragraph, the group might just ignore the paragraph and talk about something else (silent ignore), or they might send a "rolling eyes" emoji.

The Goal: Can the AI Learn?

The researchers wanted to see if the AI could do two things:

  1. Notice the "Ouch": When the group reacts negatively (sanctions), does the AI realize, "Oh, I messed up"?
  2. Change the Tune: Does the AI stop doing the thing that got the "ouch" and start acting like the rest of the group?

The Results: The "Big Kids" vs. The "Little Kids"

They tested six different AI models. Think of the results like a classroom of students taking a test on social cues:

  • The Top Performers (Claude Opus 4.7 and Gemini 3.1 Pro): These are the "big kids." When they got a negative reaction from the group, they quickly realized their mistake. About 80% of the time, they apologized or changed their behavior to fit in. They also got better the more they saw the other characters following the rule.
  • The Struggling Performers (Open-weight models like Llama, Mistral, Gemma): These models were like students who didn't quite get the hint. Even after the group rolled their eyes or ignored them, these AIs often kept doing the same thing. They didn't seem to connect the group's reaction to a need to change their behavior.

A Surprising Discovery: The "Default Personality"

The paper found something interesting about why some AIs failed. Even when the group had a specific rule (like "be very formal"), many of the AIs defaulted to being casual and chatty.

  • The Analogy: Imagine a job interview where everyone is wearing suits. If you walk in wearing a t-shirt and shorts, you break the rule. The study found that many AIs are "programmed" to be casual and friendly by default. Even when the group clearly wanted formal behavior, the AI kept acting casual, ignoring the group's signals to be serious.

How They Checked the Work

Since the "other characters" in the chat were also AI, the researchers had to make sure those characters were actually following the rules and not messing up the test. They used a "referee" (another AI) to watch the chat logs and verify:

  • Did the group actually follow the secret rule?
  • Did they actually punish the subject when it broke the rule?
  • Did the subject actually fix its behavior?

What This Means (According to the Paper)

The paper concludes that the most advanced AI models are getting much better at reading the room. They can look at how a group reacts to a mistake and adjust their behavior accordingly. However, smaller or less advanced models still struggle to pick up on these subtle social signals.

The authors emphasize that this is a specific test of social adaptation in short conversations. It doesn't mean these AIs are "smart" in a human sense or that they understand morality; it just means they are getting better at following the invisible rules of a specific group chat.

In short: SCENE is a test to see if AI can learn to "read the room" and stop making social faux pas when a group of friends gives it a subtle (or not-so-subtle) hint that it's behaving oddly. The best models are learning fast; the others are still a bit tone-deaf.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →