← Latest papers
💻 computer science

XPASS-Vis: A Dataset for Cross-Domain Personalized Image Aesthetic Assessment

This paper introduces XPASS-Vis, the first dataset designed for cross-domain personalized image aesthetic assessment featuring 6,526 stimuli across art, fashion, and landscape domains rated by 129 annotators, and establishes baseline unsupervised domain adaptation models that demonstrate the partial transferability of personalized aesthetic preferences while highlighting the need for further research to bridge the remaining performance gap.

Original authors: Takato Hayashi, Hiroaki Takahara, Candy Olivia Mawalim, Hiromi Narimatsu, Akisato Kimura, Shiro Kumano, Shogo Okada

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: Takato Hayashi, Hiroaki Takahara, Candy Olivia Mawalim, Hiromi Narimatsu, Akisato Kimura, Shiro Kumano, Shogo Okada

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a friend who is a huge fan of minimalist art. You might assume they would also love minimalist fashion or minimalist landscape photography. But would they? Maybe they love the art but think the fashion is boring, or vice versa.

This is the core puzzle the paper XPASS-VIS tries to solve: Can we teach a computer to understand your specific taste in one area (like art) and use that knowledge to guess what you'll like in a completely different area (like fashion), even if the computer has never seen your feedback on fashion before?

Here is a breakdown of the paper's journey, using simple analogies.

1. The Problem: The "One-Size-Fits-All" Trap

Currently, computers that judge beauty (Aesthetic Assessment) are like general critics. They look at a painting and say, "Most people think this is beautiful." But you and your best friend might have totally different opinions.

To fix this, researchers built "Personalized" models that learn your specific taste. However, these models usually only work in one lane. If you train a computer to know your taste in Art, it doesn't know how to apply that to Fashion or Landscapes. It's like teaching a chef to make perfect sushi, but then asking them to bake a cake without any new instructions. They might fail because they don't know how to transfer their skills.

2. The Solution: A New "Taste Test" Dataset

To study this, the authors created XPASS-Vis, the first dataset designed specifically to test this "cross-domain" transfer.

  • The Participants: They gathered 129 people (annotators).
  • The Menu: They showed these people 6,526 images across three very different "courses":
    1. Art: Paintings and sculptures.
    2. Fashion: Clothes and outfits.
    3. Landscape: Scenic views (originally video clips, turned into still photos).
  • The Rating: Each person rated every image not just on "How pretty is this?" (1 to 7), but also on specific feelings like "Does this make me feel nostalgic?" or "Is it intellectually challenging?"

The Key Feature: Unlike previous datasets where people only rated one type of image, here, the same 129 people rated all three types. This is the "Rosetta Stone" that allows researchers to see if a person's love for simple lines in art translates to a love for simple lines in fashion.

3. The Experiment: The "Unsupervised" Challenge

The researchers set up a tricky game to test if computers could learn these connections without help.

  • The Setup: They gave the computer labeled data (ratings) for Art (the Source).
  • The Challenge: They asked the computer to predict ratings for Fashion (the Target), but they gave it zero Fashion ratings to learn from. The computer had to figure out the connection on its own.
  • The Method: They tried various "Domain Adaptation" techniques. Think of these as different strategies to translate a book from one language to another without a dictionary.
    • Adversarial methods: Trying to trick the computer into forgetting which domain it's looking at.
    • Feature Alignment: Trying to mathematically line up the "shape" of Art data with the "shape" of Fashion data.

4. The Results: A Qualified "Yes"

The results were encouraging, though not perfect.

  • The Baseline: If the computer just guessed based on Art data and applied it to Fashion without any special tricks, it performed very poorly (like a bad translation).
  • The Improvement: The best "translation" methods (specifically those designed for regression tasks, called RSD and DARE-GRAM) managed to recover about 60% of the potential performance.
    • Analogy: Imagine the perfect score is 100. Without help, the computer scores 15. With the best new tricks, it jumps to 45. It's not perfect, but it's a massive leap, proving that your taste in Art does share some DNA with your taste in Fashion.

Key Finding: The computer learned that if you like "simple, clean" art, you likely lean toward "simple, clean" fashion, even without being told explicitly.

5. Who Benefits? (The Mystery)

The researchers dug deeper to see which people benefited most from this transfer. They looked at age, gender, nationality, and personality traits (like being "open to new experiences").

  • The Surprise: They found no pattern.
    • It wasn't just the "art experts" who benefited.
    • It wasn't just the "young people" or "men."
    • The benefit seemed to be random across these groups.
  • The Conclusion: The ability to transfer taste isn't about who you are on paper (your demographics); it's about the specific, invisible structure of your personal preferences. The computer can't predict who will benefit just by looking at a profile; it just works broadly for everyone to some degree.

6. Limitations: What the Paper Didn't Say

The authors are honest about what they didn't do:

  • Cultural Bias: Almost all participants were from East Asia. We don't know if this works for people from other cultures.
  • Scope: They only looked at Art, Fashion, and Landscapes. They didn't test Architecture or Product Design.
  • Video vs. Photo: The landscape data was originally video, but they used single frames. They admit they might have missed the "feeling" of movement.
  • Future Work: They didn't test "few-shot" learning (where you give the computer just 5 examples of your fashion taste) or other advanced methods. They left that for the next researchers.

Summary

The paper introduces a new dataset (XPASS-Vis) that proves personal taste is transferable. While a computer can't perfectly predict your fashion taste just by knowing your art taste, it can get about 60% of the way there using smart mathematical tricks, without needing to see your fashion ratings first. This suggests that our aesthetic souls have a consistent core that spans different types of beauty.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →