← Latest papers
💻 computer science

Where Do Images Come From? Analyzing Captions to Geographically Profile Datasets

This paper investigates the geographical bias in large-scale multimodal datasets by using LLMs to map image-caption pairs to countries, revealing that training data is heavily skewed toward high-GDP, English-speaking nations and lacks significant representation from the Global South.

Original authors: Abhipsa Basu, Yugam Bahl, Kirti Bhagat, Preethi Seshadri, R. Venkatesh Babu, Danish Pruthi

Published 2026-02-11
📖 4 min read☕ Coffee break read

Original authors: Abhipsa Basu, Yugam Bahl, Kirti Bhagat, Preethi Seshadri, R. Venkatesh Babu, Danish Pruthi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The "Global Mirror" Problem: Why AI Sees the World Through a Western Lens

Imagine you are building a giant, magical photo album that is supposed to contain a picture of everything in the world. You ask the album, "Show me a picture of a house," and it shows you a suburban home with a white picket fence in America. You ask, "Show me a road," and it shows you a wide, paved highway in Europe.

You might think, "Wait, what about the colorful houses in Morocco? What about the winding dirt paths in the Andes? What about the bustling streets of Mumbai?"

This paper, "Where Do Images Come From?", is a scientific investigation into why these "magical photo albums" (which we call AI models like Stable Diffusion) have such a skewed view of our planet.


1. The "Rich Neighbor" Bias (The Core Discovery)

The researchers looked at the massive collections of images and text used to train AI. They wanted to know: Where are these images actually from?

To do this, they used a smart AI (an LLM) to read the captions of millions of pictures. If a caption said, "A beautiful sunset in Paris," the AI knew to tag that image as France.

The Finding: The data is incredibly lopsided. It’s like throwing a massive party where 80% of the guests are from the US, UK, and Canada, while people from Africa and South America are barely invited.

The researchers found a direct link between money and visibility: the wealthier a country is (measured by GDP), the more likely its images are to be in the AI's "brain." If a country is economically powerful, the AI "sees" it; if not, the AI is essentially blind to it.

2. The "Stereotype Trap" (Diversity vs. Frequency)

You might think, "Okay, so the US has more pictures. But surely, the pictures of India or Mexico are diverse and colorful!"

The researchers tested this using a "Diversity Score." They found that just because a country has more pictures doesn't mean those pictures are varied.

The Analogy: Imagine two travelers.

  • Traveler A (Norway) takes 1,000 photos, but they are all of the same snowy mountain from slightly different angles.
  • Traveler B (Mexico) takes only 100 photos, but they include a street market, a wedding, a police car, and a mountain.

Even though Traveler A took more photos, Traveler B’s collection is much more "diverse." The researchers found that AI training data often suffers from this. It might have many images of a country, but they all look the same, which leads to stereotyping.

3. The "Uncanny Valley" of AI Generations

Finally, the researchers tested the AI itself. They asked it to generate images of specific things in specific places, like "A car in India."

The Result: The AI produced images that looked "realistic" (they didn't look like fake cartoons), but they had terrible coverage.

When asked for an Indian car, the AI tended to generate old, broken-down, or dilapidated vehicles. It captured the idea of a car in India, but it missed the reality of the modern, diverse, and shiny cars actually driving on Indian roads today. It’s like an artist who has only ever seen one postcard of a city and thinks the whole city looks exactly like that one card.


Why does this matter?

If we use these AI models to design cities, write history books, or create movies, we are accidentally teaching the world to see through a very narrow, Western-centric window.

The takeaway: To make AI truly "intelligent" and "global," we can't just scrape the internet and hope for the best. We have to intentionally go out and find the stories, the colors, and the landscapes of the entire world—not just the parts that are easiest to find online.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →