← Latest papers
💰 quantitative finance

Using Large Language Models for Idea Generation in Innovation

This study demonstrates that while AI-generated product ideas exhibit lower novelty and diversity compared to human-generated ones, they significantly outperform human ideas in average purchase intent and are seven times more likely to rank among the top 10% of high-quality concepts.

Original authors: Lennart Meincke, Karan Girotra, Gideon Nave, Christian Terwiesch, Karl T. Ulrich

Published 2026-07-31
📖 5 min read🧠 Deep dive

Original authors: Lennart Meincke, Karan Girotra, Gideon Nave, Christian Terwiesch, Karl T. Ulrich

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are standing in a vast, foggy field called "The Idea Landscape." Your goal is to find the single most valuable treasure hidden somewhere in this field. In the world of business and innovation, this treasure is a brilliant new product idea that people will love to buy. For a long time, the only way to find these treasures was to send out teams of human explorers. These humans would wander around, digging up ideas one by one. The rule of the game has always been that you don't need every single idea to be perfect; you just need to find a few absolute gems. If you dig up 100 ideas and 99 are rocks but one is a diamond, you've won. But what if you could hire a magical, super-fast robot that could dig up 1,000 ideas in the time it takes a human to dig up ten? That is the question scientists are asking about Artificial Intelligence (AI). Specifically, they are testing "Large Language Models" (LLMs)—super-smart computer programs trained on almost everything ever written on the internet. These programs are great at writing stories and answering questions, but can they be as creative as humans when it comes to inventing new things? This paper sets out to see if these digital dreamers can actually out-invent real people.

The researchers from top universities decided to put this to the test with a specific challenge: invent a new physical product for college students that costs $50 or less. They gathered three different groups of "idea generators." The first group was a team of university students who had taken a product design class before AI tools like ChatGPT were widely available. The second group was a computer program (OpenAI's GPT-4) asked to generate ideas with a simple instruction, like a blank canvas (this is called "zero-shot" prompting). The third group was the same computer program, but this time it was shown a few examples of great ideas first to help it get in the mood (this is called "few-shot" prompting).

To see who did the best, the researchers didn't just ask experts; they asked regular people. They showed hundreds of these product ideas to a large group of college-aged people and asked, "How likely are you to buy this?" They turned these answers into a "purchase score" to measure quality. They also used computer tools to check how similar the ideas were to each other and asked people to rate how "new" or "novel" the ideas felt.

Here is what they found, and it's a bit surprising. First, the AI was actually better at the main job: making ideas people wanted to buy. On average, the ideas generated by the computer were rated higher than the ideas made by the students. The computer that got a few examples to start with (few-shot) did the best of all. It's like the robot had a better "ear" for what customers wanted.

However, there was a catch. While the AI was better at making good ideas, it wasn't as good at making different ideas. The computer's ideas tended to look and sound very similar to each other. If you imagine the human students exploring the whole foggy field, the AI seemed to stay in one or two specific spots, digging up many variations of the same few treasures. The computer's ideas were also rated as slightly less "novel" or unique by the people reading them. It's as if the robot was very good at polishing existing concepts but struggled to think outside the box in the way humans sometimes do.

But here is the most exciting part. In innovation, the average idea doesn't matter as much as the best idea. The researchers looked at the top 10% of all the ideas they collected—the absolute best ones. They found that the AI was seven times more likely to produce an idea that landed in this top 10% compared to the humans. For every one great idea the students came up with, the computer came up with seven.

The authors suggest that this is actually a conservative estimate. They didn't even count how much faster the computer was. A human might take 15 minutes to come up with five ideas, while the computer could generate 200 ideas in that same time. So, the computer isn't just better at quality; it's a productivity monster.

The paper concludes that while AI might not be as wild or diverse as a human brainstorming session, it is a powerhouse for finding high-quality, marketable ideas. It suggests that in the future, companies might not need to spend so much time trying to come up with ideas from scratch. Instead, they could let the AI generate hundreds of high-quality options quickly, and then have humans focus on picking the best ones and making them even better. The robot isn't replacing the human dreamer, but it is giving them a super-powered shovel to dig up diamonds much faster.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →