CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity
This paper introduces CreativityPrism, a scalable, cross-domain evaluation framework that assesses large language models across eight tasks in divergent thinking, creative writing, and logical reasoning using three dimensions (quality, novelty, and diversity), revealing that while frontier models excel in writing and reasoning, they show no significant advantage in divergent thinking and that performance rarely generalizes across different creative dimensions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant, magical library of text generators (Large Language Models, or LLMs). Everyone is asking: "Are these machines actually creative?"
The problem is that "creativity" is a slippery concept. Sometimes it means writing a funny story, sometimes it means solving a math problem in a weird way, and sometimes it means thinking of a new use for a paperclip. Before this paper, researchers were trying to measure creativity with different rulers for different jobs, or they were asking humans to grade every single answer, which is slow and expensive.
Enter "CreativityPrism."
Think of CreativityPrism as a new, high-tech prism that you shine a beam of light (the AI's output) through. Instead of just seeing a white blob, the prism splits the light into three distinct colors, revealing the true nature of the machine's creativity.
The Three Colors of Creativity
The authors argue that to truly understand if a machine is creative, you have to measure it in three specific ways:
- Quality (The "Does it Work?" Color): This is the foundation. If a machine writes a story that makes no sense or solves a math problem incorrectly, it doesn't matter how "new" the idea is. It's just noise. Quality measures if the output actually fulfills the task.
- Novelty (The "How Weird Is It?" Color): This measures how far the answer strays from the ordinary. Did the machine come up with a solution no human has ever seen before? Or did it just copy what's in its training data?
- Diversity (The "How Many Options?" Color): This measures the breadth. If you ask the machine for 10 ideas, are they 10 totally different ideas, or are they just 10 slightly different versions of the same idea?
The Analogy: Imagine a chef.
- Quality: The food tastes good and isn't burnt.
- Novelty: The chef invents a flavor combination no one has ever tried (e.g., chocolate and pickles).
- Diversity: The chef can make 10 different types of desserts, not just 10 variations of chocolate cake.
- CreativityPrism checks all three. A chef who makes 100 different versions of a burnt chocolate cake has high diversity but low quality and low novelty.
The Big Test: 17 Chefs in the Kitchen
The authors took 17 of the smartest AI models currently available (from companies like OpenAI, Google, Anthropic, and DeepSeek) and put them through 8 different challenges across three areas:
- Divergent Thinking: Like the "Alternative Uses Test" (e.g., "What can you do with a brick other than build a wall?").
- Creative Writing: Writing stories or poems based on prompts.
- Logical Reasoning: Solving math problems or writing code under strict, weird constraints.
What They Found (The Plot Twist)
1. The "Big vs. Small" Gap
The massive, expensive AI models (the "Frontier" models) are generally better at Creative Writing and Logical Reasoning. They are like master chefs who can follow complex recipes and write beautiful menus. However, when it comes to Divergent Thinking (coming up with wild, unrelated ideas), the gap between the super-expensive models and the smaller, open-source models almost disappears. The big models aren't necessarily better at "thinking outside the box" than the smaller ones.
2. The "Specialist" Problem
Here is the most surprising finding: Being good at one type of creativity doesn't mean you are good at another.
- A model might be amazing at writing a novel (High Quality) but terrible at coming up with weird new uses for a spoon (Low Novelty).
- A model might generate 1,000 different answers (High Diversity) but none of them are actually new or useful (Low Novelty).
The authors found that "Novelty" is the most chaotic dimension. Just because a model is good at being "new" in a math problem doesn't mean it will be "new" in a story. They are totally different skills.
3. The "Judge" Solution
Since humans are too slow to grade all these tests, the authors used a clever trick: they used other AIs to grade the answers. But they didn't just trust them blindly. They had human experts grade a small sample first to make sure the "AI Judges" were actually agreeing with humans. It's like having a head chef taste a few dishes to make sure the sous-chefs are grading the food correctly before they grade the whole kitchen.
The Bottom Line
CreativityPrism is a new tool that stops us from saying "This AI is creative" as a single, vague statement. Instead, it tells us: "This AI is great at writing good stories, okay at solving math creatively, but not very good at coming up with wild, unrelated ideas."
It proves that creativity isn't one single superpower; it's a collection of different skills that don't always travel together. To build truly creative machines, we need to train them specifically for these different "colors" of creativity, not just hope they get it all right by accident.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.