COCI: Conference Organisers and Content Identifier
This paper introduces COCI, an AI-based framework that utilizes Large Language Models and semantic mapping to extract structured metadata from unstructured Calls for Papers, thereby integrating grey literature into established scholarly knowledge graphs for systematic analysis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science is often imagined as a collection of finished books and polished journal articles, the final reports of research that have passed through strict editorial gates. Yet, a vast amount of the scientific conversation happens long before those final pages are printed. This is the world of "grey literature," a term for documents that exist outside traditional publishing channels. Among these are "calls for papers," which are essentially invitations sent out by conference organizers to gather new research. These documents are vital because they reveal the earliest sparks of new ideas, showing which topics researchers are excited about and who is leading the charge. However, because these invitations are often scattered across email lists and simple web pages, they lack the neat, standardized labels that libraries and databases use to organize information. This makes it nearly impossible to study them on a large scale, leaving a significant blind spot in our understanding of how science evolves and who is shaping it.
To solve this problem, a team of researchers has built a new digital tool called COCI, which stands for Conference Organisers and Content Identifier. Think of this system as a highly skilled translator that can read messy, unstructured text and turn it into a clean, organized list of facts. The researchers fed raw text from forty different conference invitations into their system, covering fields from computer science to materials science. The tool uses advanced artificial intelligence to scan the text and pull out specific details: the name of the conference, where and when it will happen, the broad topics being discussed, and, most importantly, the names and roles of every person organizing the event.
What makes this tool particularly powerful is what it does after it extracts those names. Instead of just listing a person as "John Smith," the system works to connect that name to a permanent, unique digital ID used by the global research community. It checks its findings against several massive, established databases to ensure it has found the right person and the right organization. If the system finds a match, it attaches a unique identifier to that person's profile, linking them to their past work and their institution. It does the same for the conference itself, connecting the event to known records of similar gatherings. This process transforms a simple, informal invitation into a structured piece of data that can be linked with other scientific records, effectively bridging the gap between informal announcements and the formal maps of scientific knowledge.
The researchers tested their system by processing forty real-world examples and then manually checking the results to see how well it performed. The tool successfully extracted the metadata from every single document it analyzed. In one specific instance, the system's output revealed a mismatch in an external database; investigation showed this was not a failure of COCI's alignment algorithm, but rather an existing error within the OpenAlex database itself, which had listed the organiser's name as an alternative alias for another researcher. This instance highlights that the reliability of semantic extraction is sometimes contingent upon the quality of the external Linked Data. The team also built a visual interface that displays this information clearly, showing the conference details, the list of organizers with their verified identities, and the topics mapped to standard scientific categories. This structured representation opens up pathways to connect with external databases, allowing for the future possibility of investigating whether organisers have previously engaged in academic misconduct.
The ultimate goal of this work is to give formal recognition to the often invisible labor of organizing the scientific community. By turning these scattered, informal documents into structured data, the researchers have created a foundation for a new way to track the health and quality of academic conferences. In the future, they plan to build a system that automatically hunts down new invitations as they appear, creating a living record of how scientific communities form and change. This would allow the broader community to understand the early stages of research trends and to credit the people who build the platforms for discovery, ensuring that the work of organizing science is seen and valued just as much as the research itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.