← Latest papers
💬 NLP

GENIE: A Fine-Grained Measure for Novelty

This paper introduces GENIE, a fine-grained evaluation metric designed to measure the novelty of large language model responses along task-specific features, addressing the limitations of holistic metrics and providing insights into the effectiveness of creativity mitigation methods.

Original authors: Ramya Namuduri, Manya Wadhwa, Anshun Asher Zheng, Greg Durrett, Junyi Jessy Li

Published 2026-06-12
📖 4 min read☕ Coffee break read

Original authors: Ramya Namuduri, Manya Wadhwa, Anshun Asher Zheng, Greg Durrett, Junyi Jessy Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge at a talent show where 100 contestants are asked to tell a story about a dinosaur and a computer.

Most of the stories sound exactly the same: a T-Rex finds a laptop in a jungle, gets scared, and runs away. But then, one contestant tells a story about a T-Rex who is actually a bored office worker trapped in a prehistoric body, trying to fix a broken toaster on a boat.

The Problem with Old Judges
In the past, if you wanted to know how "creative" or "new" a story was, you used a "holistic" judge. This judge would look at the whole story and give it a single score, like "7 out of 10."

  • The Flaw: This is like a teacher giving you a single grade for a math test without telling you which problems you got right or wrong. If the dinosaur story got a "7," you wouldn't know if it was creative because of the setting (the boat), the character (the office worker), or the plot.
  • The Confusion: Sometimes, a story might look new just because the writer used fancy words (paraphrasing), but the actual ideas are boring. Old judges often get tricked by these fancy words.

The New Solution: GENIE
The authors of this paper built a new tool called GENIE (Granular Evaluation of Novel Ideas with Explainability). Think of GENIE not as a single judge, but as a team of specialized inspectors.

Instead of giving one score, GENIE breaks the story down into specific "features" or categories, like:

  • Setting: Where does it happen?
  • Plot: What actually happens?
  • Character: Who is the main person (or dinosaur)?
  • Style: How is it written?

How GENIE Works (The "Question & Answer" Game)
GENIE doesn't just guess; it plays a game of "Twenty Questions" to understand the story:

  1. The Population: First, GENIE looks at a huge pile of other stories (the "population") to see what the "standard" dinosaur stories look like.
  2. The Questions: For a new story, GENIE asks specific questions based on the features.
    • Question: "What is the setting?"
    • Answer: "A boat."
  3. The Comparison: GENIE checks the answer against the pile of other stories.
    • If 99 other stories said "Jungle," and this one says "Boat," GENIE gives a high score for Setting Novelty.
    • If 99 other stories said "The dinosaur finds a computer," and this one says the same thing, GENIE gives a low score for Plot Novelty.

Why This is Better
The paper claims that GENIE is much better at spotting the truth than the old "holistic" judges.

  • It's Honest: It can tell you exactly where the creativity is. "This story is very new in its setting, but very boring in its plot."
  • It's Not Fooled by Fancy Words: If a writer just rewrites a boring story with different synonyms (paraphrasing), the old judges might think it's new. GENIE sees that the answers to its questions are still the same, so it knows the story isn't actually new.
  • It Separates "New" from "Good": Sometimes a story is weird and new but terrible to read. Old judges often mix "creativity" with "quality." GENIE separates them, measuring only how different the ideas are, not how well they are written.

Testing the Tool
The authors tested GENIE by:

  1. Surgery: They took a boring story and surgically changed just one thing (like changing the setting from a jungle to a boat). GENIE correctly spotted that only the setting changed and gave a high score for that specific part.
  2. Paraphrasing: They took a story and rewrote it with different words but the same meaning. GENIE correctly said, "This isn't actually new," while the old judges got confused and thought it was different.

The Bottom Line
GENIE is a new way to measure creativity that doesn't just give a vague score. It acts like a microscope, showing us exactly which parts of a story are fresh and which parts are just copies of what everyone else is doing. This helps researchers understand not just if AI is creative, but where it is creative and where it needs to improve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →