Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings
This paper argues that standard evaluations of style classification in frozen vision embeddings are inflated by artist recognition rather than true stylistic understanding, demonstrating that an artist-disjoint protocol reveals significant performance drops and uneven robustness across art movements.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to teach a robot how to recognize different genres of music. You show the robot thousands of songs and ask it to sort them into "Jazz," "Rock," or "Classical." If the robot gets really good at this, you might think, "Wow, it really understands the sound of Jazz!" But here is the catch: what if the robot isn't actually listening to the saxophones or the drum beats? What if it's just memorizing that "this specific singer always sings Jazz"? If you only test the robot on songs by the same singers it has already heard, it will look like a genius. But if you give it a brand-new singer it has never met, it might fail miserably. This is the exact problem scientists face when they try to teach computers to recognize art styles. They use powerful AI models (called "frozen vision embeddings") that have seen millions of images on the internet. These models are great at spotting patterns, but researchers wondered: are they actually understanding the style of a painting, or are they just taking shortcuts by recognizing the famous artist who painted it?
This paper, titled "Style or Signature?", dives into that mystery. The author took a popular AI model (CLIP) and a collection of 320 paintings from four famous 20th-century art movements: Impressionism, Cubism, Abstract Expressionism, and Surrealism. In the standard way of testing these models, researchers mix all the paintings up randomly. This means a test painting by Salvador Dalí might be compared to other Dalí paintings the model has already seen. The author argued this is a bad test because the model might just be saying, "I know this guy, he's a Surrealist," rather than understanding what makes Surrealism unique. So, they invented a stricter test called "artist-disjoint evaluation." Imagine a game where you have to guess the genre of a song, but you are forbidden from listening to any other songs by that same singer. You have to judge the music purely on its own.
When the author ran this stricter test, the results were a shock. The model's overall accuracy dropped, but not evenly. For Impressionism and Cubism, the model barely stumbled; it could still tell the styles apart even without seeing the same artist twice. This suggests these movements have a strong, shared visual "fingerprint" that the AI actually learned. However, Surrealism fell apart completely. The model's accuracy for Surrealism dropped by a massive 20 points, from a high score down to a level where it was guessing almost randomly. The author found that for Surrealism, the model wasn't really learning the style at all; it was just memorizing specific famous painters like Dalí. When those painters were removed from the test, the model had no idea what Surrealism looked like. They tested this with four different AI models, including one that only "sees" images and never reads text, and the result was the same: the model was relying on artist signatures, not artistic styles, for Surrealism. The paper concludes that to truly know if an AI understands art, we must stop letting it take shortcuts by recognizing the artist and start testing it on styles it has never seen before.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.