Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high cost and prolonged turnaround of manual relevance labeling, which hinder large-scale online search experimentation. To overcome these limitations, the study introduces vision-language models (VLMs) into industrial search relevance evaluation for the first time, establishing an automated labeling pipeline deployed in Pinterest’s online A/B experiments. The proposed approach substantially improves evaluation efficiency and coverage, enabling more granular sampling strategies and reducing the minimum detectable effect (MDE). Empirical results demonstrate strong agreement between VLM-generated relevance judgments and human annotations, confirming the method’s capacity to support high-quality, high-sensitivity assessment of search systems at scale.
📝 Abstract
Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.
Problem

Research questions and friction points this paper is trying to address.

relevance evaluation
web-scale search
human annotation
scalability
personalized search
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Automated Relevance Evaluation
Web-Scale Search
Online A/B Testing
Minimum Detectable Effect
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Han Wang
Pinterest
A
Alex Whitworth
Pinterest
P
Pak Ming Cheung
Pinterest
Z
Zhenjie Zhang
Pinterest
K
Krishna Kamath
Pinterest
X
Xi Chen
Pinterest
R
Roberto Konow
Pinterest
K
Kurchi Subhra Hazra
Pinterest