← Latest papers
💬 NLP

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

The paper introduces IDEAFix, a controlled evaluation framework for analyzing how task formulation and defixation prompting strategies influence LLMs' divergent thinking, revealing that while structured guidance can boost solution originality, inherent output homogenization remains a persistent limitation.

Original authors: F. Carichon, S. Sharma, M. Girard, R. Rampa, G. Farnadi

Published 2026-06-02
📖 5 min read🧠 Deep dive

Original authors: F. Carichon, S. Sharma, M. Girard, R. Rampa, G. Farnadi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of very smart, very well-read robots (Large Language Models, or LLMs) and you ask them to come up with new ideas for a project. You might expect them to be like a room full of brilliant inventors, each shouting out wild, unique concepts. But often, these robots end up sounding exactly the same, repeating the same safe, boring answers. This is called "fixation"—they get stuck in a rut.

The paper IDEAFix is like a new, super-organized science lab designed to test exactly how good these robots are at thinking outside the box, and how we can shake them out of their ruts.

Here is a simple breakdown of what they did and what they found:

1. The Problem: The "Hive Mind" Effect

Think of these AI models as students who have read every book in the library. When you ask them a question, they don't "invent" new things; they remix what they've already read.

  • The Issue: If you ask 10 different AI models to "design a new chair," they might all come up with slightly different versions of the same wooden chair. They are stuck in a "hive mind," producing homogenized (identical) outputs.
  • The Goal: The researchers wanted to see if they could trick the robots into breaking out of this pattern and coming up with truly original, diverse ideas.

2. The Solution: The IDEAFix Lab

The authors built a testing framework called IDEAFix. Imagine this as a giant "Idea Obstacle Course" with three main stations:

  • Station 1: The Scenarios (The Briefs)
    They created 81 different design challenges. Some were normal (like "design a new shoe"), and some were weird or tricky (like "design a way to train a caterpillar"). This was to see if the type of task changed the robot's behavior.
  • Station 2: The "Spice" (Attributes)
    They added "flavor" to the instructions. Just like adding hot sauce or sugar to a dish, they changed the adjectives.
    • Traditional: "Design a safe shoe."
    • Surprising: "Design a dangerous shoe."
    • Negative: "Design a terrible shoe."
      They wanted to see if making the task sound weird or negative would force the robot to think differently.
  • Station 3: The Hints (Prompts)
    This is the most important part. They gave the robots different "cheat sheets" based on human creativity methods.
    • Some hints were complex, like Design Thinking or TRIZ (methods used by human engineers).
    • Some hints were simple, like "Be wild," "Be absurd," or "Don't think like a robot."
    • They also tried "AI-specific" hints that told the robot to actively avoid its usual categories.

3. The Experiment

They ran this experiment on five different famous AI models (like GPT-4, Llama, and Grok). They asked each model to generate lists of solutions for every combination of Scenario + Spice + Hint. In total, they generated over 14,000 prompts.

They then used math to measure:

  • Fluency: How many ideas did they come up with?
  • Diversity: How different were the ideas from each other?
  • Novelty: How new and surprising were the ideas compared to what the robot usually says?

4. The Results: What Worked and What Didn't?

The Good News:

  • Simple Hints Win: Surprisingly, the complex, fancy human methods (like detailed Design Thinking steps) didn't work that well. The robots didn't seem to understand the complex instructions.
  • The "Wild" Button Works: The prompts that simply told the robot to "think outside the box," "be wild," or "be unconventional" worked the best. It's like telling a dog to "be a dog" vs. giving it a complex physics lecture on how to run.
  • Negative is Good: Asking the robot to design something "bad" or "negative" actually made it come up with more unique ideas. It seems breaking the "polite" rules forces the robot to look in new places.

The Bad News:

  • The Hive Mind Persists: Even with the best hints, the robots still had a hard time breaking free completely. If you asked five different robots to solve the same problem, they still tended to come up with very similar solutions. They are stuck in a shared "semantic space" (a mental neighborhood) that is hard to leave.
  • Task Matters: The type of task changed the results. Asking for a "product" made the robots more creative but less numerous. Asking for a "procedure" made them produce more ideas, but they were all very similar to each other.

5. The Big Takeaway

The paper concludes that while we can nudge these AI robots to be a little more creative by using the right "spices" (negative words) and simple "hacks" (telling them to be wild), they are fundamentally limited. They are excellent at remixing what they already know, but they struggle to create truly transformative ideas that go beyond their training data.

IDEAFix isn't just a test; it's a toolkit. It gives researchers a controlled way to keep testing these robots, ensuring we understand exactly where their creativity ends and their "fixation" begins. It's a reminder that while AI is a powerful tool for generating ideas, it still needs human guidance to truly break the mold.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →