← Latest papers
💬 NLP

Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies

This paper presents a highly accurate automated system that extracts 3.2 million structured quantitative data points and 20 million metadata entries from 76,000 energy system studies to create a comprehensive, FAIR database and interactive dashboard for analyzing modeling assumptions, identifying gaps between academic and empirical data, and guiding future research priorities.

Original authors: Maxime Gorres, Jan Göpfert, Patrick Kuckertz, Noor Titan Putri Hartono, Heidi Heinrichs, Jochen Linßen, Iain Staffel, Jann Michael Weinand

Published 2026-07-22
📖 6 min read🧠 Deep dive

Original authors: Maxime Gorres, Jan Göpfert, Patrick Kuckertz, Noor Titan Putri Hartono, Heidi Heinrichs, Jochen Linßen, Iain Staffel, Jann Michael Weinand

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world's energy future as a giant, complex puzzle. To solve it, scientists build digital models—virtual laboratories where they test ideas like "What if we switch everything to solar?" or "How do we power a city with wind?" But these models are only as good as the numbers they are fed. If a scientist guesses the cost of a battery or how long a solar panel lasts, and that guess is wrong, the whole puzzle falls apart, leading to bad decisions for our real-world energy grid. For years, finding these numbers has been like searching for needles in a haystack of millions of research papers. Researchers had to read them one by one, manually copying down facts, a slow and error-prone process that often led to different groups doing the exact same work twice.

Now, imagine a super-smart robot librarian that can read every single energy paper ever written in seconds, understand the numbers inside, and organize them into a giant, searchable treasure chest. That is exactly what this new study has built. The researchers created a system that automatically scanned 76,000 scientific studies published since 2010. Instead of just reading the words, this system learned to spot specific facts—like the price of a wind turbine, how efficient a solar panel is, or how long a battery lasts—and pulled them out to create a massive database of 3.2 million data points. This isn't just a list; it's a map of how scientists think about energy, showing where their assumptions match reality and where they might be drifting off course. By turning a mountain of text into a clear, interactive dashboard, the team has given everyone a powerful new tool to see the big picture of our energy future without getting lost in the details.

The Robot Librarian's Big Discovery

The core of this work is a tool called Quinex, which acts like a highly trained detective for numbers. While previous attempts to gather energy data were like trying to catch butterflies with a net—slow and often missing the fast ones—Quinex uses advanced artificial intelligence to sweep through 76,000 papers and catch the numbers with incredible accuracy. The result is a database containing 3.2 million structured quantitative data points and 20 million associated metadata entries. Think of metadata as the "who, what, where, and when" tags attached to every number. For example, if the database says a wind turbine costs $2,251 per kilowatt, it also knows that this number came from a study in Mauritius in 2018. This level of detail turns a simple list of numbers into a rich, searchable story.

The researchers used this database to peek behind the curtain of energy research and found some surprising patterns. First, they looked at who is doing the research. While China produces the most papers overall, when you adjust for population size, smaller countries like Denmark and the UK are actually publishing way more energy studies per person than you'd expect. It's as if a small village is producing more inventions than a giant city, suggesting a very intense focus on energy innovation in those specific regions.

They also discovered how research priorities have shifted over time. In the early 2010s, a huge chunk of research was focused on coal, but as the world moved toward cleaner energy, that focus faded, replaced by a surge in studies about solar (photovoltaics) and wind. Interestingly, the database revealed that scientists often focus on specific technologies based on their local environment. For instance, Japan's research heavily favors solar power, mirroring the fact that they have installed far more solar panels than wind turbines. In contrast, European countries have kept a steady focus on both wind and solar.

One of the most critical findings involves the accuracy of the numbers scientists use. The team compared the costs and efficiencies found in these 76,000 papers against real-world data from the International Renewable Energy Agency (IRENA). They found that while the general trends were correct, there was a systematic "lag" and a tendency for academic papers to be a bit too optimistic or pessimistic about costs. For example, the papers suggested that the cost of solar panels was dropping, but the real-world drop was actually happening even faster. Similarly, the papers often assumed wind turbines would last longer or work more efficiently than they actually do in the real world. This suggests that while scientists are on the right track, their models sometimes rely on assumptions that are slightly out of date or too idealistic.

The database also acts as a time machine, showing how the complexity of these energy models has grown. Ten years ago, a model might have treated a whole continent like a single, blurry blob. Now, thanks to better computing power, models are zooming in, breaking continents into networks of interconnected cities and regions. They are also looking at time in much finer detail, moving from looking at energy use in broad yearly chunks to tracking it hour-by-hour to better understand how to store energy when the sun isn't shining or the wind isn't blowing.

Perhaps the most exciting part of this work is what it enables for the future. Because the database is "FAIR" (Findable, Accessible, Interoperable, and Reusable), it is available to anyone through an interactive dashboard. Researchers can filter the data to find exactly what they need, whether they are studying battery lifespans in Africa or hydrogen costs in Europe. This means less time spent digging through dusty archives and more time spent solving the actual puzzle of how to power our world. The authors emphasize that while the robot librarian is incredibly fast, human experts are still needed to double-check the work and interpret the context, ensuring that the numbers make sense in the real world.

In short, this paper doesn't just give us a bigger pile of numbers; it gives us a clearer lens through which to view the entire field of energy research. It shows us where the global conversation is heading, where the gaps in our knowledge lie, and how we can build better, more accurate models to guide our transition to a sustainable future. By automating the boring part of the job, the researchers have handed the keys to the next generation of scientists, allowing them to focus on the big, creative questions that will shape our world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →