← Latest papers
💻 computer science

AGC-Bench: Measuring Artificial General Creativity

This paper introduces AGC-Bench, a comprehensive benchmark and evaluation framework that establishes a unified measure of artificial general creativity, revealing a single creativity factor distinct from general intelligence while demonstrating that top human creators still outperform current frontier LLMs.

Original authors: Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai, Anna Attuch, Namrata Shivagunde, Swastik Roy, Rajkumar Pujari, Paul V. DiStefano, Sherin Muckatira, Claire E. Stevenson, Mikhail Gronas, Anna Rumshisky

Published 2026-07-02
📖 5 min read🧠 Deep dive

Original authors: Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai, Anna Attuch, Namrata Shivagunde, Swastik Roy, Rajkumar Pujari, Paul V. DiStefano, Sherin Muckatira, Claire E. Stevenson, Mikhail Gronas, Anna Rumshisky

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant library of tests designed to see how "creative" different people are. Some tests ask you to come up with new uses for a brick, others ask you to write a funny joke, and some ask you to solve a scientific puzzle. For a long time, researchers have argued: Is creativity one big super-power that helps you do well at all these things (like how being good at math might help you with physics), or is it a collection of separate skills where you might be a great poet but terrible at engineering?

Now, we have Artificial Intelligence (AI) models that can write stories, solve problems, and tell jokes. But we didn't have a way to measure if an AI's creativity is a general super-power or just a set of isolated tricks.

This paper introduces AGC-Bench, a massive new "creativity gym" for AI. Here is what they did and found, explained simply:

1. Building the Ultimate Creativity Gym

The researchers didn't just make one new test. Instead, they acted like librarians and detectives. They scanned over 3,000 scientific papers to find every single creativity test ever made for AI. From that mountain of papers, they picked 78 distinct tests (like a "brainstorming" test, a "storytelling" test, and a "humor" test) and built a unified system to run them all at once.

Think of it like building a single gym that has a treadmill, a weight rack, a climbing wall, and a swimming pool, all calibrated to measure the same athlete. They tested 83 different AI models on this gym.

2. The "Judge" Problem and the New Referee

When you ask an AI to grade another AI's joke or story, the AI judge can be biased—sometimes too harsh, sometimes too easy. It's like having a referee who loves one team and hates the other.

To fix this, the researchers used a clever trick called Judge Response Theory. They had three different "super-AI" referees grade every answer. Then, they used math to figure out which referee was the "strict coach" and which was the "lenient coach." They adjusted the scores so everyone was judged fairly. Finally, they trained a new, open-source AI (called AGC-Judge) to act as a perfect referee that mimics this fair panel, so anyone can use it later without needing three expensive super-AIs.

3. The Big Discovery: Is Creativity One Thing or Many?

This is the main question of the paper. They asked: If an AI is good at writing stories, is it automatically good at solving science problems?

The Answer: Yes, mostly.
When they analyzed the scores, they found that creativity in AI acts like a single, powerful engine. They call this the "C factor" (like the "g factor" for general intelligence in humans).

  • The Metaphor: Imagine creativity as a flashlight. If an AI has a bright flashlight (high "C"), it shines brightly on all the different tests, whether it's writing a poem or making a joke.
  • The Stat: This single "C factor" explained 81.5% of the differences between the models. If you know how well an AI does on one creative task, you can predict how well it will do on almost any other creative task.

However, just like a human might be a better writer than a mathematician, the top AI models still have slight preferences. Some were slightly better at humor, while others were slightly better at science ideas, but the overall "creative engine" was the same.

4. Is AI Creativity Just "Reasoning" in Disguise?

Some people thought, "Maybe AI isn't being creative; it's just really good at logic and facts."
The researchers tested this by asking the models to "be creative" versus "be logical."

  • The Result: Telling the AI to "be creative" boosted its scores much more than telling it to "use its reasoning brain." This proves the test is actually measuring creativity, not just logic.

5. Humans vs. AI: Who is More Flexible?

They compared the AI results to a group of real humans taking the same tests.

  • The Finding: Humans are very "specialized." A human who is great at writing poetry might not be great at solving engineering puzzles. Their creativity is split into different buckets.
  • The AI Difference: AI is surprisingly "general." An AI that is good at one thing is usually good at everything. It's less specialized than a human.
  • The Winner: Even though AI is more flexible, the best human in the room still beat the best AI on the hardest tasks. The top human was slightly ahead of the top AI, but the AI was catching up fast.

Summary

The paper built a giant, fair, and standardized playground to measure AI creativity. They found that AI creativity is a single, general ability (like a super-battery) that powers performance across all types of creative tasks, rather than a bunch of separate skills. While the best AI is incredibly close to the best human, the top human still holds the crown, and human creativity is more specialized (good at some things, bad at others) compared to the AI's "jack-of-all-trades" approach.

They have released all their tools, data, and the "fair referee" AI to the public so others can keep testing and improving these models.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →