CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
The paper introduces CresOWLve, a novel benchmark designed to evaluate large language models on creative problem-solving tasks grounded in real-world knowledge, revealing that while models excel at factual retrieval, they struggle significantly with the non-obvious integration of information required for creative solutions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are playing a game of "Connect the Dots," but instead of dots, you are connecting entire worlds of knowledge. That is essentially what the paper CRESOWLVE is about.
Here is the story of the paper, told simply:
1. The Problem: The "Robot Who Knows Everything But Can't Think"
For a long time, we've been testing AI (Large Language Models) like we test a student who has memorized every textbook in the library. We ask them: "Who was the first president?" or "What is 2+2?" The AI answers perfectly.
But real-life creativity isn't just about memorizing facts. It's about lateral thinking. It's like being in a room with a locked door, and instead of looking for a key, you realize the door is actually a painting, and you need to walk through the frame.
Most current AI tests are like riddles with obvious answers or fake puzzles that don't exist in the real world. They don't tell us if the AI can actually think creatively or just recall information.
2. The Solution: A New "Obstacle Course" Called CRESOWLVE
The researchers built a new test called CRESOWLVE. They didn't invent fake riddles; they went to a famous Russian TV game show called "What? Where? When?" (think of it as a high-stakes, intellectual version of Jeopardy! but with riddles that require deep cultural knowledge and "aha!" moments).
They took thousands of these real-world puzzles, cleaned them up, and translated them into English.
The Analogy:
Imagine the AI is a chef who has read every recipe book ever written.
- Old Tests: Ask the chef, "How do you boil an egg?" The chef knows this perfectly.
- CRESOWLVE: The chef is given a basket of random ingredients: a banana, a wrench, and a map of Italy. The challenge is to create a dish that represents "a journey." The chef has to realize that the banana is a boat, the wrench is a steering wheel, and the map is the ocean. That is creative problem-solving.
3. What They Found: The "Knowledge vs. Wisdom" Gap
The researchers tested the smartest AI models available today (like GPT-4, Gemini, and others) on this new course. Here is what happened:
- The Fact-Checkers: When the questions were just about facts (e.g., "Who wrote this book?"), the AIs did okay. They could pull the answer from their memory.
- The Creative Gap: When the questions required connecting two unrelated things (e.g., "How is a poet like a clock?"), the AIs stumbled badly. Their scores dropped by up to 17%.
The Metaphor:
The AI is like a person with a perfect library in their head but no map to navigate it. They have all the books (facts), but they don't know how to open a book on "History" and a book on "Music" and realize they are talking about the same thing. They get stuck trying to find a direct link, missing the creative "leap" that humans make naturally.
4. The "Thinking" Models: Slow Down to Speed Up
The paper also tested a new type of AI called "Thinking Models." These are models that are forced to pause and "think" (write out a chain of reasoning) before giving an answer.
- The Result: These models did significantly better.
- The Analogy: It's the difference between a student who blurts out the first answer they remember (Non-thinking) and a student who takes a deep breath, draws a diagram, and connects the dots before speaking (Thinking). The "Thinking" models were much better at solving the creative puzzles.
5. Why This Matters
The paper concludes that while AI is getting smarter at remembering facts, it is still struggling to be truly creative. It can retrieve the pieces of the puzzle, but it often fails to see the picture they form together.
The Takeaway:
We have built AI that is a walking encyclopedia, but we haven't quite built one that is a creative genius yet. CRESOWLVE is a new ruler to measure exactly how far we have to go to bridge that gap. It shows us that the next big leap in AI won't just be about knowing more facts; it will be about learning how to connect them in surprising, human-like ways.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.