Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
This paper reveals that while large language models (LLMs) align with human creativity evaluations on intrinsic qualities like novelty, they diverge significantly due to their reliance on narrower, model-specific standards and their insensitivity to contextual information, making the choice of a specific LLM evaluator a consequential decision for creativity assessment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are in a giant, noisy art gallery where everyone is trying to guess which paintings are "creative." Some people look only at the paint and the brushstrokes. Others look at who painted it, how much it costs, or whether it's trendy at the moment. For a long time, humans have been the only judges in this gallery. But now, a new kind of judge has arrived: the Large Language Model, or LLM. Think of an LLM as a super-smart robot that has read almost every book, article, and website ever written. Because it knows so much about what humans say, we hoped it could judge creativity just like a human does. But here's the mystery: sometimes the robot agrees with us, and sometimes it seems to be looking at a completely different painting. Scientists want to know: when does the robot think like us, and when does it get lost in its own head? This paper dives into that question to see if we can trust these digital judges to pick the winners in our creativity contests.
The researchers set out to find out exactly what rules these robot judges follow. They didn't just ask, "Is this creative?" Instead, they asked the robots to rank the importance of 26 different things humans usually look for, like "is it new?", "is it useful?", or "is it famous?". They tested six of the most popular LLMs (GPT, Grok, DeepSeek, ERNIE, Claude, and Gemini) and compared their answers to what real humans say.
Here is the big surprise they found: The robots are not all thinking alike, and they aren't thinking exactly like us either. While humans use a wide net to catch creative ideas—looking at the idea itself and the world around it (like its reputation or market trends)—the robots mostly only look at the idea itself. Specifically, the robots are obsessed with novelty. If an idea is new, weird, or unexpected, the robots love it. They also like if it's fun or artistic. But when it comes to "contextual" clues—like whether a product is from a famous brand, if it's a mass-market hit, or if it's fashionable—the robots basically ignore them. To a human, a famous brand might make an idea seem more creative; to a robot, that's just noise.
The study also discovered that not all robots are built the same. Some models, like GPT and Grok, have a "wider net." They look at more of the 26 human standards, so their judgments match human opinions pretty well. Others, like Claude and Gemini, have a "narrower net." They stick strictly to the core idea and ignore almost everything else. This means if you ask a narrow robot to judge a creative idea, it might miss the mark compared to what a human would think, simply because it's ignoring the social or market context that humans care about.
To prove this, the researchers ran two more tests. In the first, they had humans and robots judge over 1,100 ideas. They found that the robots with the "wider nets" matched human scores much better than the ones with "narrow nets." In the second test, they showed humans and robots the exact same product descriptions, but for some, they added a sentence about the product's brand or market success. The humans' ratings jumped up when they saw that extra context; they thought the product was more creative because of it. The robots? Their ratings barely changed. They were completely blind to the social clues that swayed the humans.
So, what does this mean for the future? It suggests that we can't just treat all AI judges as the same. If you want to know if an idea is truly original and fresh, a robot might be a great helper. But if you need to know if an idea will succeed in the real world, fit a brand, or get people talking, a robot might miss the point entirely. The paper suggests that choosing the right AI judge is a big deal: different models apply different rules, so they will pick different winners. We need to be careful not to assume the robot sees the whole picture, because it's often looking through a very specific, narrow lens.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.