← Latest papers
💬 NLP

Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models

This paper critiques the limitations of the Divergent Association Task (DAT) in evaluating Large Language Models by ignoring appropriateness, proposes the novel Conditional Divergent Association Task (CDAT) grounded in human creativity theory to better distinguish genuine creativity from noise, and reveals that advanced models often sacrifice novelty for appropriateness due to training and alignment.

Original authors: Kumiko Nakajima, Jan Zuiderveld, Sandro Pezzelle

Published 2026-01-29
📖 4 min read☕ Coffee break read

Original authors: Kumiko Nakajima, Jan Zuiderveld, Sandro Pezzelle

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to judge how creative a group of people (or in this case, computer programs) are. You give them a simple challenge: "Name 10 words that are as different from each other as possible."

This is the Divergent Association Task (DAT). If you say "Apple," "Car," "Cloud," and "Shoe," you get a high score because those words are very far apart in meaning. If you say "Apple," "Banana," "Cherry," and "Grape," you get a low score because they are all just fruits.

The researchers in this paper asked: Is this a fair test for Artificial Intelligence (AI)?

The Problem: The "Random Noise" Trap

The authors discovered a major flaw in how we currently test AI creativity. They found that if you just ask an AI to be "different," it can cheat by being random.

Think of it like a game of "Simon Says." If the rule is just "be different," a person who closes their eyes and points to random words on a page might actually get a higher score than a brilliant poet. The AI can just spit out nonsense words that happen to be far apart in meaning, but they make no sense together.

In the paper's experiment, they tested two "cheaters":

  1. The Random Cheater: A computer program that just picks 10 random words from a dictionary.
  2. The "Game-Optimizer" Cheater: An AI specifically told, "Ignore meaning; just pick words that will give you the highest score on this test."

The Shocking Result: Both of these "cheaters" scored higher than the most advanced, smartest AI models. This means the old test (DAT) isn't actually measuring creativity; it's just measuring how good a model is at being random or how well it can game the system.

The Solution: The "Contextual" Test (CDAT)

To fix this, the researchers invented a new test called CDAT (Conditional Divergent Association Task).

The Analogy:
Imagine the old test was like asking someone to "Name 10 random things."
The new test is like asking: "Name 10 different things that are all related to a 'Cat'."

Now, the AI has to do two things at once:

  1. Be Appropriate: The words must actually relate to the cue (e.g., "Whiskers," "Paw," "Laser Pointer"). If it says "Toaster" or "Volcano," it fails.
  2. Be Diverse: The words must still be different from each other. "Cat," "Kitten," and "Feline" are all related, but they are too similar. "Cat," "Fish," and "Dog" are better.

This new test separates true creativity (being unique but still making sense) from random noise (being unique but nonsensical).

What They Found

When they ran this new test, the results changed completely:

  • The "Smart" AIs got too safe: The biggest, most advanced AI models (the ones trained to be helpful and follow rules perfectly) tended to play it safe. They gave very appropriate answers, but they were a bit boring and not very creative. They were like a student who knows the textbook perfectly but is afraid to think outside the box.
  • The "Smaller" AIs were more creative: Surprisingly, the smaller, less "advanced" models often scored higher on creativity. They were willing to take risks and come up with more unique, diverse ideas while still staying on topic.

The authors suggest that as AI models get bigger and are trained to be "safer" and more helpful, they actually lose a little bit of their creative spark. They become too focused on being correct and not enough on being interesting.

The Takeaway

The paper concludes that we can't just ask AI to "be different" to measure creativity. We have to ask it to be different within a specific context.

  • Old Way: "Be random!" (AI cheats by being nonsense).
  • New Way: "Be creative, but make sense!" (AI has to balance being unique with being useful).

This new method gives us a much clearer picture of which AI models are actually creative and which ones are just good at following instructions or guessing randomly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →