Visual Distances among Cultural Art Corpora Align with Expert-Coded Historical-Artistic Relatedness
This study demonstrates that visual distances derived from large vision models significantly correlate with expert-coded historical-artistic relatedness across diverse cultural units and epochs, suggesting that internet-scale visual training corpora inherently encode statistical regularities reflecting art-historical relationships.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Art history has long relied on the trained eye of scholars to trace the invisible threads connecting cultures across time and space. Experts have spent centuries documenting how a style of drapery might travel from Greece to India, or how the concept of a pyramid could emerge independently in Egypt and the Americas. These connections are not always obvious; they are often buried under thousands of years of change, hidden in the subtle details of ornament, shape, and material. For a human, recognizing these deep historical links requires years of study and access to vast libraries of records. But what if a machine, trained on the entire visual history of the internet, could see these same connections without ever reading a history book? This question sits at the intersection of artificial intelligence and cultural study, asking whether the mathematical patterns hidden inside computer vision systems accidentally mirror the way human experts understand the evolution of art.
A team of researchers set out to test this idea on a global scale, moving beyond the usual focus on Western art to include traditions from five continents spanning five thousand years. They gathered a massive collection of images representing fifty-three distinct cultural units, ranging from the ancient Mediterranean to Indigenous Oceanic lineages. To create a reliable standard for comparison, they recruited ten blind experts—specialists in art history, archaeology, and museum studies—to rate how closely related every possible pair of these cultural groups was. The experts were not looking for simple visual similarity, such as two paintings both using the color blue. Instead, they were asked to judge historical relatedness based on documented evidence of cultural contact and shared visual traditions. The result was a detailed map of human judgment, a consensus on which cultures are historically linked and which are not, covering over one thousand three hundred pairs of comparisons.
The researchers then turned to six different large vision models, the kind of artificial intelligence systems that power image recognition and generation tools today. These models had been trained on billions of images from the internet, learning to compress visual information into mathematical spaces where similar images sit close together and different ones sit far apart. The team measured the distance between the fifty-three cultural groups within these digital spaces. If the machines had learned nothing about art history, the distances between these groups would be random. If they had learned to mimic human experts, the groups that the experts said were historically linked should appear closer together in the machine's mind than groups that were unrelated.
The findings revealed a striking alignment. Across all six different types of artificial intelligence models tested, the mathematical distance between cultural groups consistently matched the judgments of the human experts. When experts said two traditions were historically connected, the models placed them closer together. When experts said they were unrelated, the models placed them further apart. This pattern held true even when the researchers controlled for geography, ensuring that the connection was not simply because two cultures happened to live near each other. The signal was strong enough to be detected even between cultures separated by vast oceans, such as the ancient Greek tradition and the Buddhist art of Gandhara, a classic example of cross-cultural influence that traveled thousands of kilometers.
The study did not claim that these machines possess a conscious understanding of history or that they can replace human scholars. The researchers were careful to note that the patterns the models found likely reflect the statistical regularities present in the images they were trained on, which are themselves shaped by how museums and archives have collected and categorized art for centuries. The models did not learn art history in a classroom; they learned the visual fingerprints of cultural contact because those fingerprints are embedded in the data they consumed. However, the fact that six different architectures, trained with different goals, all arrived at the same conclusion suggests that the visual evidence of cultural transmission is robust and detectable. It indicates that the internet-scale visual record, despite its biases and gaps, contains a coherent structure that mirrors the relationships documented by centuries of human scholarship.
To ensure that the results were not an artifact of the specific images chosen, the team also ran a perceptual check with four thousand five hundred participants from twelve different countries. They asked ordinary people, with no training in art history, to distinguish between broad categories of visual styles. The participants were able to tell the categories apart, confirming that the visual differences used in the study were real and perceptible to the human eye. This added a layer of validation, showing that the patterns the machines found were not just mathematical quirks but corresponded to something visually distinct in the real world.
The research also highlighted the limitations of the data itself. The way the fifty-three cultural units were defined reflected the uneven history of art history as a discipline, with European traditions broken down into many fine categories while other regions, such as parts of Africa and the Americas, were grouped into broader, less detailed units. The researchers acknowledged that this structural bias in the data likely influenced the results, and they called for future work to involve Indigenous and non-Western collaborators to create more equitable frameworks. Nevertheless, the core finding remained stable: the visual distances calculated by these powerful artificial intelligence systems align significantly with expert-coded historical relatedness.
This work offers a new way to look at the global history of art. It suggests that the vast, messy collection of images on the internet, when processed by modern artificial intelligence, can reveal the same deep structural connections that human experts have spent centuries uncovering. It does not replace the need for human interpretation or the careful study of historical records, but it provides a powerful tool for seeing patterns across scales that would be difficult to map by hand. The study confirms that the visual language of human culture, with all its migrations, exchanges, and independent inventions, leaves a statistical signature that even a machine trained on the internet can learn to read.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.