← Latest papers
🤖 machine learning

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

This paper presents a Vision-Language Model (VLM)-based automated relevance evaluation pipeline deployed in Pinterest Search that aligns with human annotations to significantly improve measurement efficiency, expand query coverage, and reduce Minimum Detectable Effects in online A/B experiments.

Original authors: Han Wang, Alex Whitworth, Pak Ming Cheung, Zhenjie Zhang, Krishna Kamath, Xi Chen, Roberto Konow, Kurchi Subhra Hazra

Published 2026-08-04
📖 3 min read☕ Coffee break read

Original authors: Han Wang, Alex Whitworth, Pak Ming Cheung, Zhenjie Zhang, Krishna Kamath, Xi Chen, Roberto Konow, Kurchi Subhra Hazra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a massive, endless library where the shelves rearrange themselves every second based on what you're thinking. This is the world of modern web search, a field of computer science dedicated to helping people find exactly what they need among billions of digital items. For a long time, the only way to know if a librarian (or a computer algorithm) was doing a good job was to hire a team of human readers to check every single book recommendation. They would read the title, look at the picture, and decide: "Is this actually what the person asked for?" But humans get tired, it takes forever, and it costs a fortune to hire enough of them to check everything. Recently, scientists have discovered a new kind of "super-reader" called a Vision-Language Model (VLM). Think of these as AI detectives that can read text and look at pictures at the same time, understanding the story behind an image just like a human does. The big question everyone is asking is: Can these AI detectives be trusted to do the job of the human readers, and if they can, will they be fast enough to help us build better search engines?

This paper from the team at Pinterest tells the story of how they taught an AI detective to grade search results and proved it works better than the old way. They built a system where the AI looks at a user's search question, the picture of the item (called a "Pin"), and all the text describing it, then gives it a score from 1 to 5 on how relevant it is. The team found that this AI is incredibly accurate, matching human judges about 83% of the time, and almost always staying within one point of a human's score. But the real magic isn't just that the AI is good; it's that it is lightning fast. While a human team might take two days to label a batch of results, the AI does it in just two hours. Because the AI is so much faster and cheaper, the team could stop using a small, random sample of search results and start using a much smarter, organized sampling method. This change allowed them to spot tiny improvements in search quality that were previously invisible. By switching to this AI-powered system, they reduced the "Minimum Detectable Effect"—the smallest change they can reliably measure—by six times. This means they can now catch subtle, meaningful shifts in how well the search engine works, ensuring that users find exactly what they are looking for without wasting money or time on slow human checks.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →