← Latest papers
🤖 AI

Quantifying Geospatial in the Common Crawl Corpus

This work leverages Gemini 1.5 to quantify the prevalence of geospatial information in the Common Crawl corpus and estimates that 18.7% of web documents contain such data, with only minimal variations between English and non-English languages, thereby establishing a foundation for investigating geospatial biases in large language models.

Original authors: Ilya Ilyankou, Meihui Wang, Stefano Cavazzi, James Haworth

Published 2026-05-07
📖 4 min read☕ Coffee break read

Original authors: Ilya Ilyankou, Meihui Wang, Stefano Cavazzi, James Haworth

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the Internet as a vast, chaotic library containing billions of books, articles, and webpages. This library is called Common Crawl. It is the enormous pile of raw text that "super-brains" (Large Language Models or LLMs), such as those powering chatbots, consume during their "childhood" to learn how to speak and think.

Recently, scientists discovered that these super-brains are becoming surprisingly good at understanding geography. They can tell you the distance from London to Paris or describe what a city looks like. Yet one major question remained open: Where did they learn this? Did the library they studied actually contain enough maps and addresses to teach them?

This article is like a team of librarians entering this massive library to count exactly how many pages contain location information.

The Mission: Counting the "Where"

The researchers wanted to know: How many webpages in this vast library actually mention a specific location?

They defined "specific location" in two simple ways:

  1. Coordinates: The exact numbers (such as latitude and longitude) that pinpoint a spot on a map.
  2. Addresses: A street name, a city, and a number (like "Main Street 123") that one could use to send a letter or find a building.

They decided not to count vague mentions like "the Eiffel Tower" unless they were accompanied by an address, as they wanted to ensure the location was precise enough to be found on a map.

The Detective Work: A Super-Scanner in Action

The library is too large for humans to read page by page. Therefore, the researchers used a powerful AI tool called Gemini 1.5 as their "super-scanner."

Think of Gemini as a very fast, very smart intern. The researchers gave it a stack of random webpages and asked: "Hey, does this page have a street address or a sentence of map coordinates? If yes, show me an example."

Since the library is so vast, they could not check every single page. Instead, they used a statistical trick (like a pollster asking 1,000 people to guess the opinion of an entire country) to select a representative sample of about 120,000 pages from three different time periods (2019, 2021, and 2024).

The Big Revelation

After the AI scanned the pages and humans double-checked the results to ensure the AI was not "hallucinating" (making things up), they found some interesting numbers:

  • The Goldmine: About 18.7% of all webpages in the library contain some kind of specific location data. That is roughly 1 in 5 pages.
  • The Mix:
    • Some pages had only an address (16.1%).
    • Some had only coordinates (7.0%).
    • Some had both (4.3%).
  • The Language Balance: Surprisingly, it did not matter whether the page was written in English or another language. The "location density" was almost the same for both. Whether the page was in Spanish, Chinese, or English, the likelihood that it had an address or coordinates was approximately equal.
  • The "Google Maps" Effect: Many of the coordinates found were simply links to Google Maps. This means that in countries where Google Maps is not the primary map app (such as China or Russia), the number of coordinate links is lower, which could explain why some location data is slightly rarer in these languages.

What This Means for the "Super-Brains"

The study concludes that the library from which the AI learned is actually full of geography. It is not empty.

Since the AI has consumed so many pages with addresses and coordinates, it is understandable that it has learned to understand space, directions, and distances. However, the researchers also noted a "taste imbalance" in the library:

  • English and .com websites are overrepresented (they make up a huge portion of the library).
  • This means the AI may have a "bias" toward the world as seen through English-speaking, US-centric websites and may miss the nuances of other parts of the world.

The Conclusion

The researchers did not create a new map or a new app. They simply conducted a census of the Internet's raw data. They found that geography is everywhere in the data used to train AI. This explains why AI is good at geography, but it also warns us that the AI's "view" of the world is heavily shaped by which websites were most common in the library it studied.

The study ends with the suggestion that this massive collection of location data (about 46 billion pages!) could be a "goldmine" for future researchers who wish to train AI specifically for geography tasks, but that is a task for future studies, not this one.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →