← Latest papers
💻 computer science

Large language model-assisted discovery of cohorts from scientific literature

This paper presents a question-driven framework that leverages large language models to automate the extraction of cohort names from scientific literature, demonstrating its ability to identify relevant cohorts for youth aggression genetics that are often missed by traditional cohort catalogues.

Original authors: Moritz Sturm, Lisa M. Berg, Inken Berg, Harishny Sarma, Jasmin Hartmann, Denissa Girschik, Gemma Roig, Christine M. Freitag, Andreas G. Chiocchetti

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Moritz Sturm, Lisa M. Berg, Inken Berg, Harishny Sarma, Jasmin Hartmann, Denissa Girschik, Gemma Roig, Christine M. Freitag, Andreas G. Chiocchetti

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast landscape of modern science, researchers often face a problem not of finding new data, but of finding the right old data. Imagine a scientist who wants to understand why some children display aggressive behavior. To get a clear answer, they cannot rely on a single group of a few hundred people; they need to combine data from thousands of individuals across many different studies. This process, known as a multi-study analysis, requires finding specific groups of people, called cohorts, that have the right characteristics: the right age, the right medical history, and the right type of measurements, such as genetic codes or brain scans. For decades, finding these groups has been a slow, manual hunt. Scientists have had to rely on their own memory, ask colleagues for recommendations, or search through specialized lists of datasets that may not cover every possible research question. These lists are helpful, but they are often organized by the type of data they hold rather than by the specific questions scientists want to ask, leaving many useful groups hidden in plain sight.

A team of researchers from Germany has developed a new way to solve this puzzle, turning the search for these data groups into a systematic, automated process. Instead of waiting for a human to read through thousands of scientific papers one by one, they built a digital workflow that uses a large language model—a type of artificial intelligence trained on vast amounts of text—to act as a tireless research assistant. The system starts with a simple question, such as "Which studies have genetic data on aggressive children?" The computer then translates this question into hundreds of specific search terms and scours the scientific literature, pulling up thousands of article titles and summaries. The artificial intelligence reads these summaries, looking for the names of the specific groups of people mentioned in the text. It extracts these names, cleans up the duplicates, and organizes them into a list for human experts to review.

The researchers tested this system on the specific challenge of finding genetic studies related to youth aggression. They fed the system a set of search terms covering children, teenagers, aggression, and genetics. The computer generated over five thousand different search queries and retrieved more than five thousand unique scientific records. From this massive pile of text, the artificial intelligence identified nearly two hundred potential groups of people. After human experts carefully checked these candidates against strict rules—ensuring the groups were large enough, included children under eighteen, and had the necessary genetic data—the final list settled on forty-four eligible cohorts. These forty-four groups represented a combined total of nearly nine hundred thousand participants, a scale of data that would have been incredibly difficult and time-consuming to assemble by hand.

The team also wanted to know if their new method was finding things that the old, established lists missed. They compared their list of forty-four eligible groups against the results from four major, well-known databases that scientists usually rely on. The established databases, when searched with the same question, found twenty-seven of the forty-four groups. However, they completely missed seventeen of the groups that the new, literature-based system found. This gap showed that the traditional lists are not a complete map of the available data. The new method did not replace the old lists; instead, it acted as a powerful complement, uncovering valuable resources that were documented in scientific papers but not yet indexed in the central catalogs.

To ensure the artificial intelligence was doing its job correctly, the researchers compared its work against human experts. They asked two human reviewers to check a random sample of the same scientific papers to see if they contained the right groups. The reviewers did not always agree with each other, which is common when interpreting complex text, but the artificial intelligence's performance fell comfortably within the range of human agreement. It was highly accurate at avoiding false alarms, meaning it rarely claimed a paper contained a group when it did not. While it missed a few groups that the humans found, its ability to process thousands of papers in a fraction of the time made it a practical and reliable tool. The entire process, from searching to extracting names, cost very little in terms of computing resources compared to the hundreds of hours of human labor it replaced.

The study concludes that this approach offers a practical, flexible way to build a list of data groups for any specific research question. By translating a researcher's curiosity into a systematic search of the scientific record, the method turns the vast, fragmented world of published papers into a usable inventory of human data. It does not require a pre-existing list of groups to work; it finds them by reading the stories scientists have already told about their work. This capability is particularly valuable for niche research areas where no central database exists, or for questions that cut across different types of data. The researchers have made their tools and the resulting list of forty-four youth aggression cohorts available to the public, inviting others to use this method to discover data for their own fields. The work demonstrates that while curated lists of data are essential, the scientific literature itself holds a complementary treasure trove of information that can now be accessed with the help of intelligent, automated tools.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →