← Latest papers
💬 NLP

Georeferencing Non-Gazetteered Place Names using Biological Specimen Records

This study presents a method for georeferencing historical, non-gazetteered place names found in biological specimen records by leveraging spatial relations across multiple records, demonstrating that probabilistic inference outperforms large language models in achieving high spatial precision.

Original authors: Aneesha Fernando, Surangika Ranathunga, Kristin Stock, Raj Prasanna, Christopher B. Jones

Published 2026-08-10
📖 7 min read🧠 Deep dive

Original authors: Aneesha Fernando, Surangika Ranathunga, Kristin Stock, Raj Prasanna, Christopher B. Jones

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but instead of finding a missing person, you are trying to find a place that doesn't exist on any map. This is the world of georeferencing, a branch of science that turns words describing a location into actual coordinates (like latitude and longitude) that computers can use. Usually, this is easy: if a text says "Paris," a computer looks it up in a digital dictionary of places called a gazetteer and finds the exact spot. But what happens when the text mentions a place that isn't in the dictionary? Maybe it's an old name from 100 years ago, a local nickname only the neighbors know, or a specific spot like "the big oak tree near the creek" that never got an official name. These are called non-gazetteered place names. They are like ghosts in the machine—real places that the computer simply can't see because they are missing from its reference books. This matters because millions of historical records, from old letters to scientific logs, are filled with these invisible names, locking away valuable geographic knowledge that we can't access or use.

This paper dives into a treasure trove of such hidden locations: biological specimen records. Think of these as the field notes of scientists who have spent over a century collecting plants, insects, and fungi. When they found a new species, they wrote down exactly where they found it, often using local, unofficial names for landmarks. The researchers in this study asked a big question: Can we use the clues scattered across thousands of these old notes to figure out where these "ghost" places actually are? They treated the specimen records like a giant puzzle. If one note says a plant was found "north of Cabstand" and another says it was found "south of Cabstand," and we know the exact coordinates of where those plants were picked, we can work backward to find where "Cabstand" must be. By combining these clues, the team tried to pin down the locations of these missing places using three different detective strategies: a strict rule-following method, a math-based probability method, and a modern Artificial Intelligence (AI) method.

The Mystery of the Missing Places

The story begins with a massive collection of data from the Allan Herbarium in New Zealand. The researchers looked at over 13,000 digital records of plant specimens. They scanned the text for place names and then checked them against the official government map databases. As they suspected, they found hundreds of names that didn't exist in any official gazetteer. These were the Non-Gazetteered Place Names (NGPs). Some were historical farm stations, others were old trig points (survey markers), and some were just local community spots like a specific school or a tearoom.

The problem was that these names had no coordinates. To solve this, the team realized they had a secret weapon: redundancy. In the world of field biology, scientists often collect many samples in the same general area. This means the same hidden place name might appear in dozens of different notes, each time with a different description. One note might say, "Found near Cabstand," while another says, "Found 1 mile north of Cabstand." Even though "Cabstand" isn't on a map, the specimens around it are. By knowing exactly where the specimens were found, the researchers could reverse-engineer the location of the missing place name.

Three Ways to Solve the Puzzle

To crack the code, the team tested three different approaches, each with its own personality.

1. The Strict Rule-Follower (Deterministic Method)
Imagine a robot that follows instructions perfectly but gets confused if the instructions are slightly vague. This method tried to draw a strict "box" around the possible location of the missing place. If a note said "north of," the robot drew a line and said, "The place must be above this line." If another note said "south of," it drew another line. The answer had to be where all the lines crossed.

  • The Result: This method was very picky. If the clues didn't line up perfectly (which they often didn't, because old notes can be vague), the robot gave up and said, "I can't solve this." It failed to find a location for 108 out of the 365 places it tried to solve. When it did work, it wasn't very accurate.

2. The Math Wizard (Probabilistic Method)
This approach was more like a weather forecaster. Instead of drawing hard lines, it used math to calculate the likelihood of a place being in a certain spot. It treated every clue as a "vote." A note saying "1 km east" cast a strong vote for a spot 1 km east. A note saying "near" cast a softer, wider vote. The computer added up all the votes to find the spot with the highest total score.

  • The Result: This was the champion. It successfully found a location for every single place. It was the most precise, with a median error of 1.43 km. This means that half the time, the guessed location was within 1.43 km of the real spot. It also got the location within 1 km of the truth 36% of the time.

3. The AI Detective (LLM Method)
This method used a Large Language Model (an advanced AI) and gave it the raw text of the notes and the coordinates, asking it to "figure it out." The AI had to read the clues, understand the relationships (like "north of" or "near"), and do the mental math to guess the spot. The researchers even asked the AI to explain its reasoning, which surprisingly helped it get better.

  • The Result: The AI was a strong contender. It found locations for all the places and had a very low average error, suggesting it rarely made huge, crazy mistakes. However, it wasn't as precise as the Math Wizard. Its median error was 1.80 km, and it only got within 1 km of the truth 31% of the time.

What the Clues Tell Us

The study found some interesting patterns about how these methods handle different types of clues.

  • Distance matters: When a note said "1 km north," all methods did better. The distance gave a hard anchor.
  • Direction alone is tricky: If a note just said "north" without a distance, the Math Wizard struggled a bit more because "north" could mean anywhere from a few meters to hundreds of kilometers away. The AI, however, was surprisingly good at guessing the right distance even without a number, likely using its training on how people usually describe things.
  • The "Near" problem: The most common clue was just "near." This is vague. The Math Wizard handled this well by treating "near" as a fuzzy circle around the specimen. The AI also did well, but it sometimes relied on its own internal knowledge of how the world works, which is a double-edged sword (it might guess right by luck, or guess wrong based on a bad assumption).

The Verdict

The paper concludes that while AI is incredibly smart and flexible, traditional mathematical modeling (the Probabilistic Method) is still the king of precision when you need to find a specific spot on the ground. The AI is great at making a good guess and explaining its thinking, but if you need the most accurate location possible, the math-based approach wins.

The researchers are careful to say this isn't a magic solution that fixes everything. The methods work best when there are multiple clues (at least three notes mentioning the same place) and when the specimen coordinates are accurate. They also note that their test was done on a "pseudo" dataset where they knew the answers beforehand to check their work, so real-world application might have different challenges.

Ultimately, this study proves that we don't have to throw away old, messy field notes just because they mention places that aren't on modern maps. By treating these records as a network of clues, we can bring these "ghost" places back to life, filling in the gaps in our digital maps and preserving the geographic history of our world. The next time you see a note saying "found near the old mill," you'll know that with enough other notes, we can probably find exactly where that mill used to stand.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →