← Latest papers
🤖 AI

FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data

This paper introduces FashionMV, the first large-scale multi-view fashion dataset, and ProCIR, a multimodal framework that addresses the "view incompleteness" in existing Composed Image Retrieval systems by enabling product-level retrieval through mechanisms like two-stage dialogue and caption-based alignment.

Original authors: Peng Yuan, Bingyin Mei, Hui Zhang

Published 2026-04-14
📖 5 min read🧠 Deep dive

Original authors: Peng Yuan, Bingyin Mei, Hui Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are shopping for a new outfit online. You find a shirt you like, but you want to change a few things: "I love the front, but I want the back to be open with crisscross straps, and I want the sleeves to be shorter."

In the world of current AI shopping assistants, this is a frustrating experience. Here is why, and how this new paper, FashionMV, fixes it.

The Problem: The "One-Photo" Blind Spot

Right now, most AI image search tools work like a detective who only looks at one single photo of a crime scene.

In fashion e-commerce, a product is like a 3D object. To understand it fully, you need to see the front, the back, the side, and the details. But current AI systems are "blind" to the parts of the shirt you can't see in the photo you uploaded.

  • The Scenario: You upload a photo of the shirt's front. You ask for a "backless design."
  • The AI's Failure: The AI looks at the front photo, sees a solid back (because it's hidden from view), and thinks, "I can't find a backless shirt because this one has a back." It gets confused because it doesn't know that the product has a back view that isn't in your photo.

The authors call this "View Incompleteness." It's like trying to describe a whole house by only looking at the front door and ignoring the rest of the building.

The Solution: The "Product Passport" (FashionMV)

To fix this, the researchers built a massive new dataset called FashionMV.

Think of this dataset not as a pile of individual photos, but as a collection of Product Passports.

  • Instead of just one photo, every "product" in this database comes with a passport containing 2 to 5 photos (front, back, side, detail).
  • They used super-smart AI (Large Multimodal Models) to read these passports and write detailed descriptions, creating over 220,000 "search scenarios" where a user changes one product into another using text.

This is the first time anyone has taught an AI to understand that "Front View + Back View = One Complete Product."

The New AI Brain: ProCIR

They also built a new model called ProCIR (Product-level Composed Image Retrieval). Imagine this model as a fashion stylist with a unique way of thinking. Instead of just glancing at a photo, it uses three special tricks:

  1. The "Two-Stage" Conversation (The Detective's Notebook):

    • Old Way: The AI tries to look at the photo and read your text request at the exact same time. It gets confused.
    • ProCIR Way: It breaks the job into two steps.
      • Step 1: "Okay, let me look at all the photos of this shirt (front, back, side) and memorize what it is." (It creates a pure visual memory).
      • Step 2: "Now, tell me what you want to change."
    • Analogy: It's like a chef tasting the ingredients first, then listening to your order to decide how to cook it, rather than trying to taste and listen simultaneously.
  2. The "Caption Anchor" (The Label Maker):

    • The AI reads the detailed written descriptions (captions) of the clothes and forces its visual brain to match those words.
    • Analogy: It's like putting a clear, written label on every item in a warehouse so the robot knows exactly what "crisscross strap" looks like, rather than just guessing by shape.
  3. The "Thinking Aloud" (Chain of Thought):

    • Before giving the answer, the AI is trained to "think out loud" about the differences between the clothes.
    • Analogy: It's like a student solving a math problem by writing down the steps on paper, rather than just guessing the answer.

The Big Discovery

The researchers ran a huge experiment (like a science fair with 16 different versions of the AI) and found three surprising things:

  • The "Label Maker" is the MVP: The most important part of the system was simply teaching the AI to match images with their written descriptions. Without this, everything else failed.
  • Thinking Aloud isn't always needed: If you train the AI first with a "homework" phase (Supervised Fine-Tuning) where it learns the rules of fashion, it doesn't need to "think out loud" anymore. It already knows the answer. The "thinking" step becomes redundant, like a calculator that already knows the multiplication tables.
  • Small is Beautiful: Their model is tiny (0.8 billion parameters) compared to other massive AI models (which are 10x bigger), yet it beats them all. It's like a compact sports car beating a heavy truck in a race because it's built specifically for the track.

Why This Matters

This paper changes the game for online shopping.

  • For You: You can finally upload a photo of a dress you like and say, "I want the same style, but with a deep V-neck in the back and red buttons," and the AI will actually understand that the "back" exists even if your photo didn't show it.
  • For the Industry: It moves AI from "finding similar pictures" to "understanding products."

In short, FashionMV teaches AI to stop looking at just the front door and start seeing the whole house.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →