← Latest papers
🤖 machine learning

Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers

This paper reveals that contrastively pretrained vision-language models contain a substantial, invariant subspace of shared multi-modal noise that can be safely pruned without harming downstream performance, thereby offering new mechanistic insights into their latent representational structure.

Original authors: Jakub Grzywaczewski, Dawid Płudowski, Przemysław Biecek

Published 2026-05-15
📖 3 min read☕ Coffee break read

Original authors: Jakub Grzywaczewski, Dawid Płudowski, Przemysław Biecek

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot assistant (like CLIP) that has been trained to understand both pictures and words. You ask it, "Show me a picture of a dog," and it finds one instantly. It seems perfect. But this paper argues that this robot is actually carrying around a heavy, useless backpack full of junk that it doesn't need to do its job.

Here is the breakdown of what the researchers found, using simple analogies:

1. The "Static" in the Signal

Think of the robot's brain as a giant radio receiver with 768 different channels (dimensions) listening to the world.

  • The Good Channels: Most of these channels are tuned to the "signal"—the actual meaning of the picture or the word (like "dog," "church," or "happy").
  • The Bad Channels: The researchers discovered that about 164 of these channels are just broadcasting "static." This isn't random noise; it's a specific, shared kind of static that appears in every picture and every word the robot sees, regardless of what the content is.

2. The "Ghost" Pattern

The most surprising part is that this "static" is identical across different groups.

  • The Analogy: Imagine you take a photo of a Siberian Husky and a photo of a Church building. They have nothing in common. Yet, if you look at the "junk" part of the robot's memory for both images, the junk looks almost exactly the same (91% overlap).
  • The Finding: The robot has a "ghost" pattern that gets stamped onto everything it sees. It's like if a printer had a smudge on the lens; every single document it prints, whether it's a recipe or a love letter, would have that same smudge in the corner. The researchers found that this smudge is so consistent that they can map it out perfectly.

3. The "Junk Drawer" Experiment

The researchers decided to test what happens if they simply throw away this "junk drawer" (the 164 dimensions of noise).

  • The Result: They took the robot, removed those 164 channels, and asked it to do its usual tasks (like identifying animals in photos or matching words to images).
  • The Surprise: The robot didn't get worse. In fact, it got slightly better at matching words to images. It was like taking a backpack full of rocks off a runner; they didn't slow down, they actually ran faster because they weren't weighed down.

4. Why Does This Happen?

The paper suggests that this "noise" isn't a bug in the code, but a side effect of how the robot was built.

  • The Analogy: Imagine a factory that builds cars. The engineers designed the assembly line so efficiently that the cars are great, but the vibration of the machinery accidentally imprints a tiny, useless scratch on every single car door. The scratch doesn't stop the car from driving, but it's there on every single one.
  • The researchers found that this "scratch" (the noise) is a fundamental part of the architecture, not just a mistake with one specific dataset.

Summary

The paper claims that modern AI models like CLIP are carrying around a massive amount of "shared noise" (about 164 dimensions in some versions) that has nothing to do with the actual meaning of the images or text.

  • The Good News: You can cut this noise out without hurting the AI's performance.
  • The Big Picture: A huge chunk of the AI's "brain space" is actually just storing this repetitive, non-meaningful pattern, rather than useful knowledge. By cleaning it out, the AI becomes slightly more efficient and accurate.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →