An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
This paper proposes an agentic hybrid pipeline that combines top-down grounding in Wikidata with bottom-up reflexion to autonomously generate a scalable, self-healing multilingual skills knowledge graph from noisy, unstructured HR data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, chaotic library where everyone is shouting out what they are good at. Some people speak perfect English, others speak French, German, or Spanish, and some are using slang that hasn't even been invented yet. In the world of hiring, this library is a mess. If a company wants to find a "Project Manager," they might miss a brilliant candidate who wrote "Web Project Management" or "Gestion de Projet" because the computer doesn't know those phrases mean the same thing. This is where Knowledge Graphs come in. Think of a Knowledge Graph as a super-organized map that connects ideas. Instead of just a list of words, it's a web where "Project Management" is a big hub, and "Web Project Management" is a smaller branch hanging off it. The goal is to take that messy library of shouting voices and turn it into a clean, accurate map so computers can actually help humans find the right jobs.
But building this map is tricky. You can try to draw it from the top down, using a strict rulebook (like a government list of jobs), but that misses all the cool, new, weird jobs that don't exist in the rulebook yet. Or, you can try to build it from the bottom up, letting a computer guess and group things on its own, but that often leads to a chaotic mess where the computer invents fake jobs or splits one job into ten different names. This paper introduces a clever new way to build this map by mixing the best of both worlds: a strict rulebook for the known stuff and a smart, self-correcting robot team for the new stuff.
The researchers, working with a European freelancer platform called Malt, built a system that acts like a team of detective robots. They call it an "Agentic Hybrid" approach. Here is how it works: First, the system takes a messy skill written by a human, like "gestion de projet web" (French for web project management), and tries to match it to a trusted, pre-existing map called Wikidata. Think of Wikidata as the "Wikipedia of facts" that the robots trust completely. If the skill matches something in Wikidata, great! The robot locks it in place, ensuring it's accurate and consistent across all five languages they care about (English, French, German, Dutch, and Spanish).
But what if the skill is brand new and doesn't exist in Wikidata? That's where the "Agentic" part shines. Instead of just giving up or making things up, the system uses a special "reflection" loop. Imagine a robot that writes a draft, then stops and asks itself, "Wait, is this right? Does this fit with the other skills?" If the robot realizes a skill is too specific to fit under the main "Project Management" umbrella, it doesn't throw it away. Instead, it creates a new, special "orphan" branch for it. It gives this new branch a unique ID and a clear name, like "Web Project Management," and links it back to the main hub. The system then runs this process over and over, checking its own work, merging duplicates, and cleaning up the mess until the map is as accurate as possible.
The paper shows that this method works really well, but it's not magic—it's rigorous. They tested it on over 36,000 raw, messy skill descriptions from real freelancers. The system managed to organize a massive portion of them, but it also had to be honest about what it couldn't handle. Out of the total, about 19% of the attempts to match skills to the map were "wrong guesses" because the input was too confusing or ambiguous. Furthermore, the system actively rejected another 13.7% of inputs because it determined they simply weren't valid skills (like "NOT_A_SKILL" errors) or were semantic noise. However, for the data it did successfully process, the results were impressive: it cleaned up and organized 27,743 of the inputs, grouping them into about 13,300 clear, standard skill categories. This means they reduced the chaos by more than half for the valid data, turning thousands of confusing variations into a neat, structured list. Even better, the system was able to speak all five languages perfectly, creating 66,490 standardized labels so a French speaker and a German speaker could find the same job.
The researchers found that their "hybrid" approach was much better than just using a strict rulebook or just letting the computer guess. By anchoring the known skills to Wikidata, they avoided the computer making up fake jobs (a problem called "hallucination"). By using the "reflection" loop for the new skills, they didn't miss out on the cool, emerging jobs that rulebooks haven't caught yet. The result is a living, breathing map of skills that can update itself as the job market changes, making it much easier for companies to find the right people and for people to find the right work. It's like having a librarian who not only knows every book in the library but also instantly knows how to file a brand-new book that just arrived, making sure everyone can find it no matter what language they speak—while also knowing exactly when to say, "I'm not sure what this is," rather than guessing wrong.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.