← Latest papers
🤖 AI

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers

This paper systematically evaluates existing human creativity tests for predicting Large Language Model (LLM) performance across writing, divergent thinking, and scientific ideation, revealing their limited validity and introducing the Divergent Remote Association Test (DRAT) as the first robust instrument capable of predicting scientific ideation by simultaneously assessing both convergent and divergent thinking.

Original authors: Samuel Schapiro, Alexi Gladstone, Jonah Black, Heng Ji

Published 2026-05-14
📖 5 min read🧠 Deep dive

Original authors: Samuel Schapiro, Alexi Gladstone, Jonah Black, Heng Ji

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant library of very smart robots (Large Language Models, or LLMs). You want to know which ones are truly "creative." But how do you measure creativity in a machine?

For a long time, scientists have been trying to use the same tests they use on humans to grade these robots. It's like trying to measure a fish's ability to climb a tree because you want to know how good it is at climbing. This paper argues that while some of these old tests work okay, they often fail to tell the whole story, especially when it comes to scientific ideas.

Here is a simple breakdown of what the researchers found and what they built to fix it.

1. The Problem: Using the Wrong Rulers

The researchers looked at two main types of "creativity tests" that humans take:

  • Divergent Thinking: Asking someone to come up with many different, weird answers to a question (like "list 10 things that are totally different from each other").
  • Convergent Thinking: Asking someone to find the one correct answer that connects three unrelated things (like "cottage," "Swiss," and "cake" \rightarrow "cheese").

The team tested dozens of AI models using these human tests. They discovered a few surprising things:

  • The "Smartness" Trap: Many tests that seemed to measure creativity were actually just measuring how "smart" or knowledgeable the robot was. If a robot is good at math and reading, it naturally scores high on these tests, even if it isn't actually being creative. It's like grading a student on a math test and calling it a "creativity test" just because they got a high score.
  • The "Scientific" Blind Spot: The most shocking finding was that none of the existing human tests could reliably predict if an AI could come up with good scientific ideas. You could have a robot that is great at writing stories or listing random words, but terrible at inventing a new scientific theory. The old rulers just didn't work for this job.
  • The Paradox of Newer Models: As AI models got newer and more powerful, they actually got worse at divergent thinking (coming up with unique, weird ideas). They became more rigid and standard, like a factory producing the same cookie over and over, rather than a chef inventing a new recipe.

2. The Solution: A New Hybrid Test (DRAT)

Since the old tests failed at predicting scientific creativity, the authors invented a new test called the Divergent Remote Association Test (DRAT).

Think of the old tests as two separate tools:

  • Tool A (Divergent): A sledgehammer that smashes things apart to see how many different pieces you get.
  • Tool B (Convergent): A screwdriver that tightens things together to find the one perfect fit.

The problem with scientific creativity is that you need both at the same time. You need to smash ideas apart to find something new, but then you need to screw them together to make sure they actually make sense and fit the problem.

How DRAT works:
Imagine you give the AI three very different "anchor" words (like "heartbeat," "pipeline," and "topology").

  1. The Challenge: The AI must generate 10 words that are all different from each other (Divergent).
  2. The Catch: But, every single one of those 10 words must also be able to metaphorically connect to all three anchor words (Convergent).

It's like asking a poet to write 10 completely different poems, but every poem must secretly be about a specific, complex machine. If the AI can do this, it proves it can think wildly and logically at the same time.

3. The Results

When they tested this new DRAT tool:

  • It worked! It was the first test that could successfully predict which AI models were good at coming up with scientific ideas.
  • It wasn't just a sum of parts: The researchers tried to combine the old "smash" test and the old "screw" test to see if that would work. It didn't. The new DRAT tool did something special that the two old tools couldn't do on their own. It proved that to measure scientific creativity, you have to test the ability to be wild and logical simultaneously.

The Big Takeaway

The paper concludes that we can't just use old human tests to grade AI creativity.

  • If you want to know if an AI is good at writing stories, use the "Divergent Association Task" (the weird word list).
  • If you want to know if an AI is good at divergent thinking (listing many options), use the "Conditional Divergent Association Task."
  • But if you want to know if an AI can invent new science, you need the new DRAT test, because it checks if the machine can be both wildly imaginative and logically connected at the same time.

In short: To measure a robot's scientific genius, you need a test that forces it to juggle fire (divergence) while walking a tightrope (convergence). The old tests only asked it to do one or the other.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →