← Latest papers
💬 NLP

Anticipating Innovation Using Large Language Models

This paper introduces TechToken, a transformer-based model that analyzes collective shifts in patent language to detect early signals of forthcoming technological combinations decades in advance, thereby improving the forecasting of innovation and outperforming state-of-the-art models in patent-related tasks.

Original authors: Enrico Maria Fenoaltea, Filippo Santoro, Giordano De Marzo, Segun Taofeek Aroyehun, Andrea Tacchella

Published 2026-05-07
📖 4 min read☕ Coffee break read

Original authors: Enrico Maria Fenoaltea, Filippo Santoro, Giordano De Marzo, Segun Taofeek Aroyehun, Andrea Tacchella

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to predict a surprise party. You can't know exactly who will show up or what gift they'll bring until the moment it happens. But, if you listen closely to the conversations happening in the weeks before the party, you might notice a pattern: people start talking about the same themes, using similar jokes, and planning around the same ideas. Even though no one has officially sent an invitation yet, the language of the group has already shifted to prepare for the event.

This paper argues that innovation works the same way.

The Big Idea: Innovation is a "Language Shift"

Usually, we think of a new invention (like a smartphone combining a camera and a phone) as a sudden "Eureka!" moment. The authors suggest that before two technologies ever get combined in a patent, the way inventors talk about them starts to merge.

Think of technologies as different dialects. For years, "battery technology" and "touchscreen technology" might have been spoken in completely different rooms by different groups of people. But, years before they actually get glued together in a new product, the inventors in both rooms start using similar words, metaphors, and sentence structures. The "adjacent possible" (the next big thing) leaves a whisper in the collective language of patents long before it becomes a shout.

The Problem: Old Maps vs. New GPS

Previous methods tried to predict these combinations by looking at a simple map: "How often do these two things appear together?" It's like trying to predict a storm by only counting raindrops. It's too slow and misses the subtle shifts in the wind.

Other methods tried to read the text of patents, but they treated the technical codes (like the IPC codes, which are like library catalog numbers for inventions) as static labels. They didn't understand that the meaning of a code changes depending on the story it's told in.

The Solution: TechToken (The "Translator" Model)

The authors built a new AI tool called TechToken. Here is how it works, using a simple analogy:

Imagine a library where every book has a special tag (the IPC code).

  • Old Way: You take all the books with a "Battery" tag, read them, and make one giant, blurry summary of what "Battery" means. You lose the nuance.
  • TechToken Way: The AI learns that the "Battery" tag is actually a word in its own vocabulary, just like "apple" or "run." It reads the story around the tag. It learns that in a story about electric cars, "Battery" means one thing, but in a story about pacemakers, it means something slightly different.

By teaching the AI to treat these technical codes as words in a sentence, it can understand the context of the invention. It learns that when the language around "Batteries" starts sounding more like the language around "Solar Panels," a combination is coming.

What They Found

  1. The Signal is Early: The authors found that this "linguistic convergence" happens decades before the actual invention. They could see the signal rising up to 20 years before two technologies were officially combined in a patent.
  2. It's a Group Effort: You can't find this signal in a single patent or by listening to one inventor. It only appears when you listen to the entire chorus of thousands of inventors. It's a collective shift in how the world thinks about technology.
  3. It Works Better: When they tested TechToken against other massive AI models (including huge ones like LLaMA), TechToken won. It was able to predict new combinations much more accurately, even though it was a smaller, more efficient model. It proved that understanding the context of the technical codes is more important than just having a huge database of facts.

Why This Matters

The paper concludes that the "adjacent possible" isn't silent. It leaves a readable trace in the way we write and talk about our inventions. By listening to the collective language of inventors, we can see the future taking shape long before the first patent is filed.

In short: Innovation isn't just about building new things; it's about the world starting to speak the same language about those things long before they are built. TechToken is the tool that finally learned to listen to that conversation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →