← Latest papers
💬 NLP

Breaking the HISCO Barrier: Automatic Occupational Standardization with OccCANINE

This paper introduces OccCANINE, an open-source tool fine-tuned on 15.8 million multilingual data pairs that automates the mapping of occupational descriptions to HISCO and other classification systems with 96% accuracy, thereby replacing slow manual coding and democratizing access to high-quality occupational data for economic research.

Original authors: Christian Møller Dahl, Torben Johansen, Christian Vedel

Published 2026-02-26
📖 4 min read☕ Coffee break read

Original authors: Christian Møller Dahl, Torben Johansen, Christian Vedel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a historian trying to understand the lives of people from 200 years ago. You have millions of old documents—census records, marriage certificates, ship logs—filled with handwritten notes about what people did for a living. One says "fishing and farm work," another says "servant," and a third says "weaver."

To study this data, you need to turn these messy, handwritten descriptions into neat, standardized categories (like "Farmer" or "Fisherman") so a computer can count and analyze them. This process is called occupational coding.

The Problem: The "HISCO Barrier"

For decades, researchers have used a standard system called HISCO to do this. But there's a massive problem: it's incredibly slow and boring.

Imagine you have a stack of 10,000 job descriptions. If you are a super-fast human coder, it takes you 10 seconds to read one and type in the right code. That's 28 hours of non-stop work just for a tiny sample. If you have a million records, you'd need a team of people working for years. Plus, humans get tired, make typos, and interpret things differently. One person might think "servant" is a "domestic worker," while another thinks it's "unskilled labor." This inconsistency creates a "barrier" that stops research from happening.

The Solution: OccCANINE (The "Smart Translator")

The authors of this paper built a tool called OccCANINE to break this barrier. Think of OccCANINE as a super-smart, tireless translator that has read millions of job descriptions and knows exactly how to categorize them.

Here is how it works, using a simple analogy:

1. The "Brain" (The Model)

Imagine a student who has spent their entire life reading 15.8 million examples of job descriptions paired with their correct codes. They have read records in 13 different languages (English, Danish, German, Swedish, etc.).

  • Old Way: A computer program tries to match words using a "find and replace" rule (e.g., if it sees "farm," it guesses "Farmer"). This fails if the word is misspelled ("frm") or if the description is complex ("fishing and farming").
  • OccCANINE Way: The model understands meaning, not just words. It knows that "he fishes and tends to his farm" is actually two jobs: Fisherman and Farmer. It can handle typos, old-fashioned spelling, and messy handwriting because it "understands" the concept of the job, not just the letters.

2. The Two "Modes" of Operation

The tool offers two ways to work, like two different types of assistants:

  • The "Fast" Mode (Flat Decoder): This is like a speed-reader. It scans the text and instantly spits out a list of likely codes. It's incredibly fast (minutes for thousands of records) but requires you to set a "confidence filter" (e.g., "only accept answers I'm 90% sure of").
  • The "Good" Mode (Sequential Decoder): This is like a careful detective. It builds the answer one digit at a time (e.g., "6... 1... 1... 1... 0"). It's slightly slower but more robust, especially for complex descriptions where a person had multiple jobs. It doesn't need you to set filters; it just gives you the best answer.

3. The "Magic" Result

The results are staggering:

  • Speed: What used to take weeks of human labor now takes minutes.
  • Accuracy: It gets the answer right 96% of the time. That's as good as, or better than, a human expert.
  • Versatility: It doesn't just work for HISCO. The authors showed it can be easily retrained to work on other historical coding systems (like US or British census codes) with very little extra work.

Why This Changes Everything

Think of OccCANINE as a key that unlocks a giant library.
Before, researchers could only read a few books because translating the titles took too long. Now, they can read the whole library in a day.

  • Democratization: You don't need a team of expensive research assistants anymore. A single researcher with a laptop can process massive datasets.
  • New Discoveries: Because the data is now easy to analyze, historians can ask new questions. They can track how women's jobs changed over 200 years, how the Industrial Revolution shifted the workforce, or how geography influenced careers, all with high-quality data.
  • Reliability: Since the computer is consistent, every researcher gets the same result for the same data. No more "human error" or "interpretation bias."

The Bottom Line

OccCANINE is a tool that turns the impossible task of manually sorting millions of old job descriptions into a simple, automated process. It replaces the "bottleneck" of human labor with a smart, fast, and accurate AI, allowing historians and economists to finally see the full picture of how people worked throughout history.

It's not just a tool; it's a time machine that lets us process the past as fast as we can imagine the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →