A Pathway for Assessing Grey Literature: Leveraging AI to Extract Conference Metadata and Organiser Information from Calls for Papers
This paper introduces COCI, an AI-based framework that leverages Large Language Models to automate the extraction and structuring of granular metadata from unstructured Calls for Papers, thereby enabling systematic metascience analysis of previously overlooked grey literature.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science is often imagined as a collection of finished books, neatly stacked on library shelves, where researchers study the past by reading what has already been published. But a vast amount of important knowledge exists outside these formal channels. This is called "grey literature," a term for reports, working papers, and documents that never make it into standard journals. Among the most vital of these are "calls for papers," which are announcements inviting researchers to submit their work to upcoming conferences. These documents are the early warning system of science, revealing new trends and emerging fields before they are ever written up in a final article. However, because these announcements are scattered across the web in messy, unstandardized formats, they have been nearly impossible to study on a large scale. Traditional computer tools struggle to read them, leaving a huge gap in our understanding of how scientific communities actually form and evolve.
To bridge this gap, a team of researchers has developed a new system called COCI, which uses advanced artificial intelligence to read and organize these chaotic announcements. The system acts as a digital archivist, taking the raw, unstructured text of a conference invitation and turning it into a clean, organized list of facts. It identifies the name of the event, the year and location, and, most importantly, the people running the show. It does not just list names; it figures out who these people are, what institutions they belong to, and what specific roles they play, such as organizing a workshop or managing publicity. By doing this, the system transforms a jumble of text into a structured database that scientists can actually use to analyze the landscape of research.
The researchers tested this system by feeding it forty different conference announcements from a wide variety of fields, ranging from computer science and engineering to dentistry and archaeology. The goal was to see if the system could handle the messy reality of the internet, where one conference might list its organizers in a simple paragraph and another might bury the information in a complex table. The system successfully extracted the key details from all forty documents. It identified the conference series, the specific edition, and the geographic location. It also pulled out the names of every person on the organizing committee and linked them to public scientific records to verify their identities and affiliations. This process allows the system to distinguish between different people who might share the same name and to confirm which university or organization they currently represent.
One of the most significant achievements of the system is its ability to make sense of the topics these conferences cover. When a conference announcement lists its areas of interest, it often uses informal language or unique phrases. The system translates these into standard scientific concepts, creating a bridge between the casual language of the announcement and the formal vocabulary used by researchers worldwide. It also matches the conference name to existing global databases, ensuring that the event is correctly identified even if the name is slightly different or if the conference has been held under various titles over the years. This creates a reliable map of the event within the broader world of science.
However, the researchers were careful to note that the system is not perfect and requires human oversight to catch its mistakes. In one test case, the system incorrectly assumed that every member of an organizing committee belonged to the same university because the original text did not specify their individual affiliations. The researchers built a safety check into the system to catch this kind of error, recognizing that it is highly unlikely for an entire committee to come from a single institution. In another instance, the system matched a researcher to the wrong person in a public database because the database itself contained conflicting information. These examples show that while the tool is powerful, it still depends on the quality of the data it accesses and the logic of the rules it follows.
The ultimate goal of this work is to shift the focus of scientific analysis toward these non-traditional events. By making it possible to study calls for papers systematically, the system opens the door to understanding how research communities are built from the ground up. It allows for the tracking of emerging trends before they appear in formal journals and provides a way to recognize the contributions of the many people who organize and shape these events but often go unnoticed in traditional metrics. The researchers have made their tool available to the public as an open-source project, inviting others to use it and improve it. This approach suggests a future where the full, diverse picture of scientific activity can be seen, not just the polished final results, but the dynamic, evolving process of discovery itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.