← Latest papers
💻 computer science

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo

This paper evaluates how open-weight language models navigate the competing demands of lexical constraints and communicative effectiveness in the game of Taboo, revealing that while intervention strategies affect rule compliance and description quality, models still struggle significantly with constrained lexical grounding compared to humans.

Original authors: Sara Candussio, Francesca Padovani, Daniel Scalena, Malvina Nissim

Published 2026-07-02
📖 4 min read☕ Coffee break read

Original authors: Sara Candussio, Francesca Padovani, Daniel Scalena, Malvina Nissim

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a party game called Taboo. You have a card with a secret word at the top (like "Pizza") and a list of five words you are not allowed to say (like "Cheese," "Italy," "Dough," "Slice," or "Delivery"). Your goal is to describe "Pizza" to your friends so they can guess it, without using any of those forbidden words.

This paper is like a scientific experiment where researchers taught two different AI "players" how to play this game. They wanted to see:

  1. Can the AI follow the rules (not say the forbidden words)?
  2. Can the AI still describe the word well enough for someone to guess it?
  3. How does the AI compare to a human player?

Here is a breakdown of what they did and what they found, using some simple metaphors.

The Three Ways They Tried to Stop the AI from Cheating

The researchers tried three different methods to force the AI to obey the rules, kind of like three different ways to stop a child from eating a cookie before dinner:

  1. The "Ask Nicely" Method (Prompting):
    They simply told the AI, "Please describe this word, but do not use these specific words." It's like telling a child, "Please don't eat the cookie."

    • Result: The AI was pretty good at listening, but sometimes it still slipped up and used a forbidden word or a word that sounded very similar (like saying "Cheesy" when "Cheese" was banned).
  2. The "Digital Handcuffs" Method (Constrained Generation):
    They programmed the AI so that the moment it tried to type a forbidden letter or word, the system physically blocked it. It's like putting a lock on the cookie jar so the child literally cannot open it.

    • Result: This worked almost perfectly. The AI never said the forbidden words. However, because it was so scared of breaking the rules, it sometimes became too vague, making it harder for the guesser to figure out the answer.
  3. The "Brain Surgery" Method (SAEs/Internal Manipulation):
    This was the most complex. Instead of blocking words, they tried to tweak the AI's internal "thoughts" (its neural network) to make the concept of the forbidden word disappear from its mind while it was thinking. It's like trying to convince the child, "Actually, cookies don't exist right now," so they don't even think about them.

    • Result: This was a mixed bag. The AI still managed to say some forbidden words (the "surgery" wasn't perfect), but the descriptions it gave were actually very creative and easy to guess! It seems that when you just block the words, the AI gets clumsy, but when you tweak its "mind," it finds clever workarounds.

The Big Surprise: The AI is Terrible at Guessing

The researchers also flipped the game. They had the AI try to guess the word based on descriptions written by humans or other AIs.

  • The Human Advantage: When humans played, they were great at guessing. They could look at a description like "It's round, red, and you put it on pasta" and instantly think, "Tomato!"
  • The AI Struggle: The AI was surprisingly bad at this. Even when given a perfect description, the AI often couldn't guess the word. It was like having a student who can memorize a textbook perfectly but fails to understand the story when someone tells it to them.
    • The paper notes that the AI needed to make 5 to 10 guesses just to get as good at guessing as a human who gets it on the first try.

The Main Takeaway

The paper concludes that for AI, following rules and being a good communicator are two different skills that often fight each other.

  • If you force the AI to be strict about the rules, it becomes a boring, vague describer.
  • If you let it be creative, it might accidentally break the rules.
  • And no matter how good it is at describing, it is still very bad at the "guessing" part of the game.

The researchers suggest that this happens because humans have a special, intuitive way of connecting words in our brains (like a web of associations) that current AI models just don't have yet. The AI can follow instructions, but it hasn't quite learned the "human intuition" needed to play Taboo perfectly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →