← Latest papers
💬 NLP

ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution

This paper introduces ICE-ID, a novel benchmark dataset of nearly one million records from 220 years of Icelandic census data designed to address unique challenges in longitudinal identity resolution, such as patronymic naming and temporal drift, while providing comprehensive analysis artifacts and a deployment-faithful evaluation protocol.

Original authors: Gonçalo Hora de Carvalho, Lazar S. Popov, Sander Kaatee, Mário S. Correia, Kristinn R. Thórisson, Tangrui Li, Pétur Húni Björnsson, Eiríkur Smári Sigurðarson, Jilles S. Dibangoye

Published 2026-02-25
📖 4 min read☕ Coffee break read

Original authors: Gonçalo Hora de Carvalho, Lazar S. Popov, Sander Kaatee, Mário S. Correia, Kristinn R. Thórisson, Tangrui Li, Pétur Húni Björnsson, Eiríkur Smári Sigurðarson, Jilles S. Dibangoye

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to solve a massive, centuries-old puzzle where the pieces are people, but the picture keeps changing. The names on the pieces get misspelled, the borders of the countries they live in shift, and sometimes, the pieces are missing entirely.

This is the challenge researchers face when trying to track individuals through history using old census records. To help solve this, a team of scientists from Iceland and beyond has created a new, massive "training ground" for computers called ICE-ID.

Here is a simple breakdown of what this paper is about, using some everyday analogies.

1. The Problem: The "Ghost" in the Machine

Imagine you are trying to find your great-great-grandfather in a stack of 200-year-old documents.

  • The Name Game: In Iceland, last names aren't fixed; they are based on the father's name (e.g., "Jón's son" becomes Jónsson). This means thousands of people might all be named "Jón Jónsson." It's like trying to find a specific "John Smith" in a city where everyone is named "John Smith."
  • The Moving Target: People move, get married, have children, and die. The borders of towns and counties change over time. A person might be listed as living in "Farm A" in 1800 and "Farm B" in 1850, even if it's the same family.
  • The Missing Pieces: Old records are messy. Sometimes the birth year is missing, or the father's name is blank.

Most computer programs used to match records (like finding duplicate products on Amazon) are trained on clean, static data. They fail miserably when faced with this messy, moving, historical reality.

2. The Solution: The "Time-Traveling Detective" Dataset

The authors built ICE-ID, a dataset containing nearly 1 million records from 16 different census waves in Iceland, spanning from 1703 to 1920.

Think of this dataset as a gym for AI detectives.

  • The Workout: Instead of just matching "Shoe A" to "Shoe B," the AI has to match "Jón Jónsson, age 30, living in Farm X in 1820" to "Jón Jónsson, age 55, living in Farm Y in 1845."
  • The Cheat Sheet: The researchers didn't just dump the raw data; they spent years curating it. They created a "Gold Standard" list of who is who. They know for a fact that Record A and Record B belong to the same real-life person. This allows them to test if the AI is actually learning or just guessing.

3. Why This is Different (The "Special Sauce")

Most existing datasets for testing AI are like static snapshots. They are like a photo of a grocery store aisle. They are easy because the products don't move, and the names don't change.

ICE-ID is different because:

  • It's a Movie, Not a Photo: It tracks people over 220 years. The AI has to deal with "time drift"—how a person's name, address, and family change over decades.
  • It's a Family Tree, Not a List: It includes sparse family links (who is the father? who is the spouse?). It's like giving the detective a few clues about the family tree to help solve the case, even though many clues are missing.
  • It's Hierarchical: It understands geography like a Russian nesting doll: Farm → Parish → District → County. If a town changes its name, the AI needs to understand that the "Farm" is still the same place, just under a new label.

4. The Challenge: The "Name Collision"

The paper highlights a funny but difficult problem: Name Collisions.
Because of the naming tradition, the most common name in the dataset, "Jón Jónsson," appears 15,599 times.

  • The Analogy: Imagine a high school where 15,000 students are all named "John." If you try to find "John" in the yearbook, you need more than just the name. You need his birth year, his parents' names, and where he lived. ICE-ID forces AI to learn how to use those extra clues to tell the "Jóns" apart.

5. The Goal: Teaching AI to be a Historian

The ultimate goal of releasing this dataset is to create better AI tools that can:

  • Track Social Mobility: See how families moved from poor farms to cities over generations.
  • Study Disease: Understand how epidemics spread through specific families or regions over time.
  • Fix History: Automatically link broken records so historians can see the full life story of an ancestor without spending 100 hours in a library.

Summary

In short, ICE-ID is a massive, open-source library of Icelandic history designed to teach computers how to be historical detectives. It moves beyond simple "matching" to teach AI how to handle the messy, changing, and confusing reality of human lives over centuries. By releasing this data, the authors hope to unlock new ways to understand our past using modern technology.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →