← Latest papers
💬 NLP

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning

This study demonstrates that multimodal large language models can effectively evaluate visual creativity in a zero-shot setting with substantial alignment to human ratings, while their step-by-step reasoning provides interpretable insights into the evaluation process despite not improving scoring accuracy.

Original authors: William Orwig, Roger E. Beaty

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: William Orwig, Roger E. Beaty

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, super-smart art critic who has never taken a class on art, never seen a human grade a drawing, and has never been told what "creativity" looks like. You just hand them a picture and ask, "On a scale of 1 to 5, how creative is this?"

That is essentially what this paper does. The researchers asked six different advanced AI models (think of them as different "digital judges") to do exactly that. They wanted to see if these AIs could guess how humans would rate the creativity of pictures, without any special training or practice.

Here is the story of what they found, broken down into simple parts:

1. The "Zero-Shot" Magic Trick

Usually, to teach an AI to grade art, you have to show it thousands of examples of "good" and "bad" drawings along with human scores. It's like teaching a student by showing them a stack of graded exams.

But these researchers tried something different: Zero-Shot. They gave the AI just the picture and a simple instruction: "Rate this 1 to 5." No examples, no training. It was like asking a stranger on the street to grade a painting they've never seen before.

The Result: Surprisingly, the AI judges were pretty good at it. Their ratings matched human ratings quite well.

  • For AI-generated images (polished pictures made by computers), the AI judges and humans agreed very strongly.
  • For hand-drawn sketches (rough, messy drawings by people), the agreement was still there, but a bit weaker.

2. The "Polished vs. Rough" Bias

The researchers noticed the AI judges had a funny personality quirk, like a critic who loves high-definition photos but hates pencil sketches.

  • The "Polished" Bias: When looking at slick, computer-generated images, the AI judges were too generous. They gave higher scores than humans did.
  • The "Rough" Bias: When looking at simple, hand-drawn sketches, the AI judges were too harsh. They gave lower scores than humans did.

It seems the AI models secretly love "elaborate" and "complex" images. If a drawing has a lot of lines and detail, the AI thinks it's creative, even if the human judge thinks it's just "busy." The AI models seemed to confuse "detailed" with "creative."

3. The "Thought Process" (The Chain of Thought)

Some of these AI models have a special feature: before they give a final score, they write out their "thought process." It's like a student showing their work on a math test.

The researchers peeked inside these thought processes to see how the AI decided on a score. They found the AI follows a consistent four-step recipe:

  1. Perception: "I see a cat and a tree." (Describing what is there).
  2. Originality: "This is weird and new!" (Judging if it's unique).
  3. Quality: "The lines are smooth and pretty." (Judging how well it's drawn).
  4. Justification: "So, I give it a 4." (Picking the number).

The Big Surprise: The researchers thought that if the AI took the time to "think" and write out these steps, it would get better at matching human scores. It didn't. The "thinking" didn't make the scores more accurate. However, it did give the researchers a window into why the AI gave the score it did. It showed that the AI was paying attention to "quality" (how pretty it is) rather than just "originality" (how weird it is), which explains why it liked the polished AI images so much.

4. Does the AI Need to Know What It's Looking At?

The researchers wondered: "If the AI gets confused and thinks a drawing of a dog is a cat, does its creativity score go down?"

The Answer: Not really.
Even when the AI misidentified what the drawing was, or when the drawing was so abstract that humans couldn't agree on what it was, the AI's creativity score still matched human scores pretty well. This suggests the AI isn't just checking a checklist of "Is this a dog? Yes/No." It's looking at the vibe, the composition, and the weirdness of the image, even if it doesn't know exactly what the object is.

The Takeaway

This paper shows that we can use these big, general AI models as "creativity judges" without needing to train them first. They are surprisingly good at it, but they have a bias: they love complex, detailed images and dislike simple, rough sketches.

By looking at their "thought processes," we can see exactly where they are getting it right and where they are getting it wrong. It's like having a critic who can explain their reasoning, even if their taste is a little different from ours. The researchers even built a free app so anyone can try this "AI judge" on their own pictures.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →