Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
This paper introduces SNIC, a human-validated testbed for evaluating Large Language Models on norm-based reference resolution, revealing that even state-of-the-art models struggle to consistently identify and apply implicit social norms essential for situated, embodied interactions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are in a busy kitchen with a friend. You ask, "Can you pass me a mug?" There are three mugs on the counter: one is sparkling clean, and two are covered in dried coffee stains.
If you are a human, you instantly know your friend will hand you the clean mug. You don't need them to say "the clean one" because you both share an unspoken rule: Don't drink from someone else's dirty cup. This invisible rule is a social norm.
This paper, titled "Where Norms and References Collide," asks a simple but tricky question: Can modern AI (specifically Large Language Models or LLMs) figure out which mug you mean without being told explicitly?
Here is the breakdown of their findings, using some everyday analogies:
The Problem: The AI's "Blind Spot"
Think of an LLM as a very smart student who has read almost every book in the library. They are great at answering trivia and solving logic puzzles. However, this student has never actually lived in a kitchen, a library, or a restaurant. They know what a "mug" is, but they haven't learned the subtle, unwritten rules of how people behave in those places.
The researchers call this missing skill Norm-Based Reference Resolution (NBRR). It's the ability to look at a messy scene, understand the "social grammar" (the rules of behavior), and guess what someone is pointing at.
The Experiment: Building a "Norm Trap"
To test this, the team created a dataset called SNIC (Situated Norms in Context).
- The Setup: They created 9,000 little stories (vignettes) involving everyday tasks like cleaning, serving food, or tidying up.
- The Trap: In every story, they asked a question that seemed ambiguous. For example, "Pick up the plate."
- Scenario A: The room is being cleaned. The AI should pick up the dirty plate.
- Scenario B: The room is being served. The AI should pick up the clean plate.
- The Human Test: Before trusting the AI, they asked 210 real humans to solve these puzzles. The humans mostly agreed on the "normative" answer (e.g., "Pick up the dirty one because we are cleaning"), proving the rules were clear to people.
The Results: The AI Gets Lost Without a Map
The researchers tested several top-tier AI models (like GPT-4, Llama, and Phi) on these 9,000 stories. The results were like watching a GPS try to navigate a city it has never visited:
- The "Guessing Game" Failure: When the AI was just given the story and asked to pick the right object, it struggled. It got the right answer less than half the time on average. It often couldn't tell the difference between a "cleaning task" and a "serving task" because it didn't understand the norm behind the task.
- The "Formal Language" Dead End: The researchers tried giving the AI a super-precise, mathematical description of the scene (using a code language called Prolog) to remove all ambiguity. Surprisingly, this didn't help much. It's like giving a driver a perfect map but telling them, "You still don't know the traffic laws." The AI knew the facts but still didn't know the rules of the road.
- The "Cheat Sheet" Success: When the researchers explicitly gave the AI a list of the rules to follow (e.g., "Rule: In a cleaning task, pick up dirty items"), the AI's performance skyrocketed. It suddenly got the answers right.
The Big Takeaway
The paper concludes that current AI models are like excellent encyclopedias but poor neighbors.
- They know what a mug is.
- They know what cleaning is.
- But they don't inherently know how people use mugs in a social context unless you explicitly tell them the rule.
The authors argue that this is a "blind spot." AI models are trained on text, but social norms are often implicit (never written down) and physically grounded (tied to real-world objects and actions). Because these rules aren't always stated clearly in books, the AI misses them.
Why This Matters (According to the Paper)
The paper suggests that if we want robots or AI assistants to work in real homes or offices, they need to learn these invisible social rules. Currently, they might hand you a dirty mug when you ask for a drink, not because they are "stupid," but because they haven't learned the unspoken social contract that says, "We don't do that."
In short: The paper built a test to see if AI understands the "unwritten rules" of daily life. It found that while AI is smart, it is currently terrible at guessing what we mean when we rely on those unwritten rules, unless we spell the rules out for it first.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.