JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes
The paper introduces JobHop v2, a publicly released dataset of over 355,000 career trajectories derived from pseudonymized multilingual resumes using a robust LLM extraction pipeline, offering rich annotations and improved reliability to advance workforce planning and labor market research.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of work as a giant, messy library where every single book is a person's resume. For years, researchers trying to understand how people move from one job to another were stuck because these "books" were written in different languages, had weird formatting, and were full of scribbles that computers couldn't read. They were like trying to build a map of a city using only torn-up, handwritten notes.
Previously, the only maps available were either tiny (like a few thousand notes), made up of fake data, or locked away in private vaults that no one else could open. There was one earlier attempt called "JobHop v1," but it was a bit glitchy; sometimes the computer got confused by the messy notes and spat out broken files it couldn't use.
Enter JobHop v2, a brand-new, super-smart librarian built using a massive Artificial Intelligence (AI) brain. The researchers gave this AI a stack of about 440,000 real, anonymized resumes from the Flemish Public Employment Service (VDAB). These resumes were a chaotic mix of Dutch, French, and English, with names and locations hidden behind <MASK> tokens to protect privacy.
The Magic Trick: The Reasoning Robot
Instead of just guessing, this new AI uses a "reasoning-controlled" approach. Think of it like a detective who doesn't just read a clue but pauses to think, "Wait, does this date make sense? Is this a job or a school class?" If the AI gets stuck or produces a messy file (like a JSON that won't open), it has a special "retry" button. It tries again with a bit more focus.
The result? Out of the roughly 400,000 resumes that made it past the initial cleaning, the AI successfully turned 100% of them into neat, organized digital files. That's a perfect score! It extracted 355,315 unique career stories, detailing 1,993,291 job experiences and 923,981 education entries.
What's in the Box?
JobHop v2 is like a treasure chest of career data. It doesn't just list job titles; it organizes them using a standard global dictionary called ESCO. It tells you exactly when someone started and stopped a job (down to the quarter of the year, like "Q1 2020"), what level of education they have (from Primary school up to a PhD), and even what specific tools they used (like Python or Spark).
Did it actually work better?
The researchers didn't just say "it's good"; they put it to the test. They compared the AI's work against three different sets of human and AI-generated "gold standard" answers.
- The Score: The best version of their AI (using a huge 120-billion-parameter model) got a score that was incredibly close to the "ceiling" of what humans and other AIs could agree on. It was only 1.1 to 2.7 percentage points away from the perfect agreement limit.
- The Showdown: When they did a blind taste-test against the old JobHop v1 (where a judge didn't know which was which), the new JobHop v2 won 68.3% of the time, while the old version only won 29.9%. The judge noted that v2 was much better at handling the messy, hidden parts of the resumes and understanding dates in different languages.
What It's NOT
It's important to know what this dataset doesn't do. The researchers explicitly ruled out using fake, made-up resumes or pre-standardized codes that were synthesized by an AI before the extraction happened. They wanted real, raw text from real people. Also, they were very careful not to guess too much: about 6.9% of the job entries were labeled "unknown" because the AI decided it was safer to admit it didn't know than to guess wrong. They didn't try to force a label on everything.
The Bottom Line
JobHop v2 suggests that we can finally turn the chaotic, unstructured mess of millions of resumes into a clean, massive dataset that helps us understand how careers actually evolve. It's not a magic crystal ball that predicts the future perfectly, but it provides the most accurate, large-scale map of career paths ever built from real, unstructured text. The researchers are releasing this map and the tools used to build it to the public, hoping others will use it to study labor markets and job recommendations without having to start from scratch.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.