← Latest papers
🤖 AI

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

This paper introduces NICE, a theory-grounded diagnostic benchmark comprising 137 items across 11 dimensions that evaluates the social intelligence of large language models, revealing that while they achieve high aggregate accuracy, they exhibit consistent weaknesses in specific communication facets such as multi-turn interaction, nonverbal cues, and synchrony.

Original authors: Yunjin Qi, Zhaojun Jiang, Xuan Wu, Hanxi Pan, Yixuan Wang, Yanfang Liu, Xiang Ji, Churu Yu, Chunyuan Zheng, Yingze Chen, Jie He, Liuqing Chen, Zaifeng Gao

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Yunjin Qi, Zhaojun Jiang, Xuan Wu, Hanxi Pan, Yixuan Wang, Yanfang Liu, Xiang Ji, Churu Yu, Chunyuan Zheng, Yingze Chen, Jie He, Liuqing Chen, Zaifeng Gao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help you navigate the complex world of human relationships. You don't just want to know if they can answer questions correctly; you want to know if they can "read the room," understand unspoken rules, and know when to speak up or stay quiet.

This paper introduces NICE (Norm, Interaction, Cognition, Experience), a new "report card" designed specifically to test how well Large Language Models (LLMs)—the brains behind AI chatbots—handle these social situations.

Here is the breakdown of what the researchers did and what they found, using simple analogies:

1. The Problem: The "Blind" Test

Before NICE, testing AI's social skills was like taking a driving test where you only had to parallel park.

  • Old Tests: Some tests only checked if the AI knew a specific fact (like "What is a polite greeting?"). Others put the AI in a chaotic, open-ended conversation where it was impossible to tell why it failed.
  • The Issue: If an AI got a low score, researchers couldn't tell if it was bad at understanding emotions, bad at following rules, or just bad at talking. It was like saying, "You failed the driving test," without knowing if you couldn't see, couldn't steer, or didn't know the traffic laws.

2. The Solution: The "Social X-Ray" (NICE)

The researchers built NICE to be a diagnostic tool, not just a scorecard. Think of it as an X-ray that breaks down "Social Intelligence" into specific parts so doctors (researchers) can see exactly where the bone is broken.

They created a framework based on psychology, organizing social skills into 4 main categories and 11 specific dimensions:

  • Norms: Knowing the rules of the game (politeness, ethics).
  • Interaction: How you talk and act with others.
  • Cognition: Understanding what others are thinking or feeling.
  • Experience: Learning from past interactions.

Instead of asking the AI to write a long story, NICE gives it a short scenario (like waiting for a crowded elevator) and asks it to rank three possible responses from "Best" to "Worst." This forces the AI to show it understands the boundaries of social behavior, not just the "correct" answer.

3. The Experiment: AI vs. Humans

The researchers tested 5 of the smartest AI models available (like GPT-5.5 and Claude-Opus-4.7) against a group of real humans.

  • The Surprise: Overall, the AI models actually scored higher than the humans! They were better at remembering facts, following ethical rules, and managing relationships in a theoretical sense.
  • The Catch: When the researchers looked at the "X-ray" (the specific dimensions), they found a massive, consistent weakness.

4. The Big Discovery: The "Communication Gap"

While the AI was great at the rules (Norms) and the logic (Cognition), it struggled terribly with Communication (D3).

Think of it like a student who memorized the entire textbook on "How to Be a Friend" but fails when actually talking to a friend.

  • Where AI Failed: The models were specifically bad at:
    • Multi-turn communication: Keeping a conversation going naturally over several turns.
    • Nonverbal communication: Understanding tone, body language, or implied meaning (even in text).
    • Synchrony: Matching the rhythm and flow of a human interaction.

The "Bow" Example:
The paper gives a funny example: In a scenario where someone bows too deeply (180 degrees), humans immediately recognized this as weird and inappropriate. However, the top AI models thought this was actually a good or neutral response. The AI was so focused on being "polite" that it missed the social boundary that said, "This is too much!"

5. The Conclusion

NICE proves that even the smartest AI models have a specific "blind spot." They are excellent at knowing what to say based on rules, but they are often clumsy at how to say it in a way that feels natural and human.

The paper concludes that to make AI truly socially intelligent, we can't just make them smarter overall; we need to specifically train them to fix these communication gaps, just like a coach would focus on a player's weak foot rather than just telling them to "play better."

In short: AI is a brilliant student of social theory, but it's still a bit awkward at the actual party. NICE is the tool that finally told us exactly why.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →