← Latest papers
💻 computer science

Using LLMs for Ontology generation by video summarization

This paper presents a two-phase algorithm that leverages Large Language Models to automatically generate RDF ontologies from YouTube video URLs by first creating a textual summary and then converting it into a structured domain representation, achieving 90% accuracy on a sports dataset.

Original authors: Khaled Omar

Published 2026-09-25
📖 5 min read🧠 Deep dive

Original authors: Khaled Omar

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast landscape of digital information, there is a persistent challenge: how to teach computers not just to store data, but to understand the relationships between ideas. Imagine a library where every book is filed by its cover color rather than its subject; finding a specific story would be nearly impossible. In the world of computing, this problem is solved by creating "ontologies." Think of an ontology as a structured map of knowledge for a specific field, such as sports or medicine. It defines the key concepts within that field and, crucially, draws the lines connecting them. This map allows machines to reason, search, and share information in a way that mimics human understanding. However, building these maps has traditionally been a slow, expensive process requiring teams of human experts to manually define every term and connection. As the volume of digital content explodes, particularly in video formats, the old methods of manual construction are struggling to keep pace.

A new approach, detailed in recent research by Dr. Khaled Omar of Damascus University, attempts to bypass the bottleneck of human labor by using advanced artificial intelligence to build these knowledge maps directly from video. The study focuses on a specific type of artificial intelligence known as large language models—systems trained on massive amounts of text that can read, understand, and generate human language with remarkable fluency. While these models have previously been used to summarize text, this research asks a bolder question: can they watch a video, understand its core message, and automatically construct a formal knowledge map of the subject matter? The answer, according to the study, is a promising yes. The researchers developed a two-step system that first turns a video into a concise written summary and then transforms that summary into a structured data format that computers can use to reason about the video's content.

The process begins with a simple input: a link to a video on YouTube. The system does not watch the video in the way a human does, by processing moving images. Instead, it first extracts the spoken words from the video, creating a text transcript. This transcript is then fed into a powerful artificial intelligence model called Llama 3.2, which the researchers ran on their own local computers to avoid cloud costs. The model acts as a skilled editor, reading the transcript and breaking it down into manageable chunks. It summarizes each section and then combines those summaries into a single, coherent narrative. This narrative is not just a random collection of sentences; it is carefully structured to highlight the main theme of the video, identify the most important people or objects mentioned, and extract the key conclusions. The result is a clean, text-based story that captures the essence of the original video.

Once this textual summary is ready, the system moves to its second phase: turning the story into a formal map. Here, a different artificial intelligence model, Gemini Pro 002, takes over. The researchers provide this model with a specific set of instructions, or a prompt, telling it to act as a knowledge architect. The model is asked to look at the summary, identify the domain of the video, and list the key concepts found within it. It then takes these concepts and begins to draw the connections between them, defining how they relate to one another. Finally, the model outputs this structure in a standard format known as RDF, which is a way of writing data that allows different computer systems to exchange and understand the information seamlessly. The entire workflow, from a raw video link to a structured knowledge map, happens automatically, requiring no human intervention to define the terms or relationships.

To test if this automated method actually works, the researchers applied it to a dataset of one hundred videos focused on sports. They did not just let the system run and hope for the best; they measured its performance against a high standard. For the first part of the process, the video summarization, they compared the AI's written summary to the video's original title to see how well they matched in meaning. The system achieved a high level of agreement, with the summaries aligning with the titles about ninety-five percent of the time. For the second part, the creation of the knowledge map, the researchers compared the AI-generated map against a map created by a human expert. This is the true test of whether the machine understands the structure of the knowledge, not just the words. In this comparison, the automated system successfully recreated the human expert's map with an accuracy rate of eighty-five percent. When looking at the entire process as a whole, the researchers calculated an overall accuracy of ninety percent.

The implications of this work are significant for how we manage the growing ocean of video content on the internet. By automating the creation of these knowledge maps, the system reduces the heavy reliance on human experts, making it possible to organize and understand video data at a scale that was previously impossible. The researchers note that while they tested this on sports videos, the method is not limited to that field; it could be applied to medical lectures, educational tutorials, or any domain where video is a primary source of information. The study suggests that the future of organizing digital knowledge may lie in these hybrid systems, where artificial intelligence handles the heavy lifting of extraction and structuring, allowing humans to focus on higher-level validation and application. The research stands as a demonstration that with the right tools, the complex task of turning a moving image into a structured, machine-readable understanding of the world is no longer just a theoretical possibility, but a practical reality.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →