Ethology of Latent Spaces
This study employs an ethological perspective to demonstrate that latent spaces in vision-language models are not neutral but exhibit distinct, training-data-driven "algorithmic scopic regimes" that generate divergent and often contradictory political and cultural interpretations of art, necessitating a critical reassessment of how these models are used in digital art history.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have three different art critics. They are all looking at the same collection of 301 paintings and sculptures, ranging from the 15th century to the 1900s. You might expect them to agree on what they see. But in this study, the "critics" aren't human; they are three different AI models (OpenAI CLIP, OpenCLIP/LAION, and SigLIP).
The paper argues that these AI models are not neutral, blank slates. Instead, each one has developed its own unique "personality" and way of seeing the world, shaped entirely by the specific data it was fed and the way it was built. The author calls this an "ethology of latent spaces"—basically, studying the "behavior" of these AI brains just like a biologist studies the behavior of animals in the wild.
Here is a breakdown of the paper's main points using simple analogies:
1. The "Invisible Glasses" Analogy
Think of these AI models as wearing invisible glasses.
- The Math: Inside the computer, the AI sees images as a cloud of numbers (vectors) in a giant, multi-dimensional space.
- The Meaning: When we ask the AI, "Is this image political?" it doesn't "think" about politics. Instead, it looks at its cloud of numbers and sees which "political" numbers are closest to the image.
- The Twist: The paper shows that the "glasses" each model wears are different. One model might see a black-and-white photo as "dark and gloomy," while another sees it as "high contrast and dramatic," simply because of how they were trained.
2. Three Different "Personalities"
The study compares three specific models and finds they act like three very different people:
- OpenAI CLIP (The Museum Curator): This model was trained on a carefully curated, private dataset. It acts like a traditional art historian. It tends to see "politics" only in famous, institutional art movements (like Cubism or Dada). If you show it an African mask, it sees it as "art" or "culture," but not necessarily as "political."
- OpenCLIP/LAION (The Internet Scroller): This model was trained on a massive, messy dump of data scraped from the entire internet. It acts like a chaotic internet user. It sees "politics" everywhere, but often in a random, noisy way. It might label a painting as "colonial" just because the word "colonial" appeared in the caption of a similar image on the web, even if the painting itself isn't about colonialism.
- SigLIP (The Critical Theorist): This model uses a different technical architecture. It acts like a sharp, modern critic. It sees "politics" in almost everything. For example, it labels African masks as highly "political" (59% of the time), whereas the other models say they are apolitical. It seems to have learned that non-Western art is inherently tied to political struggle, while Western art is "neutral."
3. The "Neon Sign" Confusion
The paper gives a great example of how these AIs get confused by the difference between a physical object and a word.
- The Test: The researchers asked the models to rank images by "brightness."
- The Result: OpenAI CLIP ranked Joseph Kosuth's neon art (which is often photographed against a black background) as one of the brightest things in the collection.
- Why? The AI didn't "see" the pixels. It saw the word "neon" in the training data and associated the word "neon" with the concept of "bright light." It confused the label with the reality.
- SigLIP, however, actually looked at the pixels and correctly identified that the dark photos were dark.
4. "Emergent Bias" vs. "Statistical Bias"
Usually, when we talk about AI bias, we think of "statistical bias" (e.g., "The dataset didn't have enough photos of women").
- The Paper's New Idea: This study introduces "Emergent Bias." This is a bias that appears out of nowhere, like a chemical reaction. It's not in the raw data directly; it emerges from the complex way the AI connects the dots between words and images.
- The Analogy: Imagine a recipe. If you leave out salt, the dish is bland (statistical bias). But if you mix flour, sugar, and yeast in a specific way, you get bread rising (emergent property). The AI's "political" views are like the rising bread—they weren't put in the bowl, but they rose up because of how the ingredients (data) were mixed.
5. The "Three Scopes" (How They Look)
The author uses a metaphor from art history to describe how these models "look" at the world:
- The Entropic Gaze (LAION): A chaotic, noisy look that picks up everything the internet says, often reinforcing stereotypes found in web searches.
- The Institutional Gaze (OpenAI): A controlled, museum-like look that reinforces established, Western art history rules.
- The Semiotic Gaze (SigLIP): A theoretical look that tries to find deep, critical meanings in everything, sometimes forcing political interpretations where humans might just see a shape.
The Bottom Line
The paper concludes that AI is not a neutral tool for art history.
If you use an AI to analyze art, you aren't just getting a "computer's opinion." You are getting the specific, hidden "personality" of that model's training data.
- If you use Model A, you might see art as "neutral and aesthetic."
- If you use Model B, you might see the same art as "deeply political and colonial."
The author warns that we cannot just "delegating" the interpretation of art to AI without understanding that the AI has its own invisible rules and biases that we can't see unless we compare different models against each other. The "truth" of what the AI sees depends entirely on which "glasses" it is wearing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.