← Latest papers
💬 NLP

Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings

This paper introduces TaxonomyBuilder, a framework demonstrating that filtering job posting data before processing yields clearer and more domain-specific AI skill taxonomies than using unfiltered corpora with LLM-enhanced clustering tools.

Original authors: Stephen Meisenbacher, Peter Norlander

Published 2026-05-21
📖 4 min read☕ Coffee break read

Original authors: Stephen Meisenbacher, Peter Norlander

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to organize a massive, chaotic library containing millions of books about Artificial Intelligence (AI). Your goal is to create a clear, easy-to-read map (a taxonomy) that helps people find exactly what they need: specific AI skills and tasks.

The authors of this paper, Stephen Meisenbacher and Peter Norlander, built a new tool called TAXONOMYBUILDER to do this job automatically. They tested it using two giant piles of real-world data: millions of job postings from the US labor market.

Here is the story of what they found, explained simply:

The Problem: The "More is Better" Trap

Usually, when people try to organize a huge mess, their instinct is to throw everything into the mix. They think, "If I include more job postings, my map will be more complete." They might even try to copy-paste similar sentences to make the pile even bigger (a process called data augmentation).

The authors asked: Is this actually helpful? Or does it just make the library messier?

The Experiment: Cooking with Different Ingredients

To find the answer, they ran 24 different experiments using their TAXONOMYBUILDER tool. They changed three main "ingredients" in their recipe:

  1. Data Augmentation (Adding more): Did they add extra, similar sentences to the mix to make the dataset bigger?
  2. Filtering (Removing the noise): Did they throw away the "weakest" or most vague job descriptions, keeping only the top 25%, 50%, or 75% of the clearest ones?
  3. Soft Clustering (Fuzzy grouping): Did they force "noisy" items that didn't quite fit into a group to join anyway, just to be safe?

The Big Surprise: Less is More

The results were counter-intuitive. The team discovered that adding more data actually made the map worse.

  • The "Augmentation" Mistake: When they added extra data to make the pile bigger, the resulting map became confusing and less accurate. It was like trying to organize a library by adding more books that were just slightly different copies of the same story; it just created clutter.
  • The "Filtering" Win: The best maps were created when they were strict. They threw away the vague, low-quality job descriptions (keeping only the top 75% or 50% of the best data). By removing the "noise," the AI could see the clear patterns much better.
  • The "Soft Clustering" Nuance: Forcing messy items into groups only helped if they hadn't added extra data in the first place. If they added extra data, forcing messy items in made things worse.

The Analogy: The Gold Panning Test

Imagine you are panning for gold (AI skills) in a massive river (the job postings).

  • The Old Way: You scoop up every bucket of water, mud, and rocks you can find, hoping to get more gold. You end up with a huge, muddy pile where it's impossible to see the gold.
  • The New Way (TAXONOMYBUILDER's finding): You carefully sift through the water first. You throw away the muddy, useless buckets (filtering). You only keep the buckets that look promising. Then, you pan those. You end up with a smaller pile, but it is pure gold.

The Conclusion

The paper concludes that when building a map of complex skills from massive amounts of text, quality beats quantity.

  • Don't try to include every single piece of data you can find.
  • Do be selective. Filter out the weak or noisy information.
  • Do let the AI focus on the clear, high-quality examples to build a map that is actually useful and easy to understand.

In short: If you want a clear map of the AI job market, you don't need to read every single job posting in existence. You just need to read the best ones carefully.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →