The Human Creativity Benchmark
The paper introduces the Human Creativity Benchmark (HCB), which distinguishes between professional convergence on technical correctness and divergence in aesthetic taste to argue that evaluating creative AI requires preserving these distinct signals rather than collapsing them into a single quality metric.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Why "Good" is Complicated
Imagine you are trying to judge a cooking competition.
- The "Objective" Judge: This judge checks if the food is actually cooked, if the salt is there, and if the plate isn't broken. Everyone agrees on this. If the steak is raw, it's a fail.
- The "Subjective" Judge: This judge asks, "Do I like the flavor?" One person might love spicy food, while another hates it. Neither is wrong; they just have different tastes.
The Problem: Most AI tests today act like they only have the "Objective" judge. They try to force everyone to agree on a single score (e.g., "This AI is 8/10"). The authors of this paper say that in creative work (like making ads, websites, or videos), this approach is broken. It throws away the most important part of the data: the fact that people should disagree on taste.
The Solution: The "Martini" of Creativity
The authors propose a new way to test AI called the Human Creativity Benchmark (HCB). They visualize the creative process as a sideways martini glass (see Figure 1 in the paper):
- The Wide Part (Ideation): At the start, you want lots of wild, different ideas. This is where "divergence" (different opinions) is good.
- The Narrowing Part (Mockup): You start picking the best ideas and making them look real.
- The Stem (Refinement): At the end, you are polishing the final product. Here, you need "convergence" (everyone agreeing) on technical details like spelling, layout, and whether the product actually works.
How They Tested It
Instead of asking, "Is this AI good?", they asked professional creatives (designers, video editors, coders) to evaluate AI outputs in three specific ways:
- Prompt Adherence: Did the AI do exactly what it was told? (Like a recipe follower).
- Usability: Could this actually be used in a real job? (Is the text readable? Is the button clickable?).
- Visual Appeal: Is it pretty? (This is where taste comes in).
They ran this test across 15,000 judgments from experts in five different fields: Landing Pages, Desktop Apps, Ad Images, Brand Images, and Product Videos.
What They Found (The "Aha!" Moments)
1. Agreement vs. Disagreement are different signals
- Convergence (Agreement): When the AI makes a mistake (like unreadable text or a broken layout), all the experts agree it's bad. When the AI follows the rules perfectly, they also agree it's good.
- Divergence (Disagreement): When the AI is technically correct, experts start to disagree based on taste. One expert might say, "This ad looks like luxury fashion," while another says, "This looks cheap." The paper argues this disagreement isn't a mistake; it's a feature. It means the AI is offering different creative directions.
2. No single AI wins everything
If you just look at an average score, you might think one AI is the "best." But the paper found that no model is the best at everything.
- Some models are great at the Ideation phase (coming up with wild, creative concepts).
- Other models are better at the Refinement phase (fixing typos and making things look professional).
- Analogy: It's like saying a Formula 1 car is "better" than a pickup truck. The F1 car wins on the race track (Refinement), but the pickup truck wins at the construction site (Ideation/Utility). You need the right tool for the specific stage of the job.
3. The "Single Score" Trap
The paper warns that if you collapse all these different opinions into one single number (e.g., "Model X is 4.2 stars"), you lose the most useful information. You stop seeing where the model is reliable (technical correctness) and where it is flexible (creative style).
The Takeaway for Humans and AI
The authors suggest that we shouldn't try to make AI agree on a single definition of "beautiful." Instead, we should:
- Use AI models that are reliable when we need technical correctness (convergence).
- Use AI models that are steerable when we need creative variety (divergence).
In short: Don't ask AI to be perfect at everything. Ask it to be perfect at the right part of the process. The benchmark helps us figure out which AI is the right "chef" for the specific course of the meal we are cooking.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.