Efficient GPU Retrieval for Semantic Search

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为提高LinkedIn语义搜索相关性,提出一种与策略对齐的检索框架,并通过两阶段GPU架构实现高效检索。
📝 Abstract
Semantic Search on LinkedIn must retrieve relevant profiles from a corpus of hundreds of millions in response to natural-language queries such as "a fintech founder in Berlin who worked in payments." The deployed relevance policy is bottleneck-oriented: every active non-negotiable facet must be satisfied, and a pre-existing LLM Graded Relevance (GR) judge operationalizes this through a fixed min/median aggregation over facet grades. Cosine similarity instead averages evidence, letting a strong match on one facet mask failure on another, capping the recall of the first-stage (L0) retriever. We present a policy-aligned retrieval framework: embeddings are partitioned into eight category-supervised segments whose scores follow the same min/median rule at serving time; for multi-vector retrieval, this segment score is computed independently per tagged document slot and maximized across slots. A lightweight single-slot Stage-1 scorer generates high-recall candidates, while scale-invariant relative-norm gating keeps category activation consistent across training, evaluation, and serving. On 21K held-out queries, this representation improves offline relevance over a matched-capacity baseline, with gains broadly distributed across facet combinations. We serve this framework with a two-stage GPU architecture: an FP8 coarse ranker scores the full corpus, increasing per-shard capacity by 71% and Stage-1 matmul throughput by 36%, then an FP16 stage exactly re-ranks an oversampled candidate set, recovering 99.6-99.8% of full-FP16 recall at over 500 QPS per shard replica. In a member-randomized A/B test, exploratory-query Precision@10 under the unchanged GR judge rises from 63.7% to 79.0% and navigational Precision@1 from 65.5% to 74.7%, with a blinded human evaluation independently confirming the Precision@10 gain.
Problem

Research questions and friction points this paper is trying to address.

Semantic Search
Relevance Policy
Cosine Similarity
Retrieval Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

policy-aligned retrieval framework
segmented embeddings
scale-invariant relative-norm gating
two-stage GPU architecture
FP8 coarse ranker
🔎 Similar Papers
No similar papers found.
D
Dhritiman Das
LinkedIn, Mountain View, CA, USA
Chujie Zheng
Chujie Zheng
Qwen Team, Alibaba Group
Artifical IntelligenceLarge Language Models
R
Ronak Kaoshik
LinkedIn, Mountain View, CA, USA
P
Pratik Dixit
LinkedIn, Mountain View, CA, USA
V
Vishal Shah
LinkedIn, Mountain View, CA, USA
Yanbo Li
Yanbo Li
LinkedIn, Mountain View, CA, USA
Jiahao Xu
Jiahao Xu
Nanyang Technological University
LLM Efficient ReasoningNMTAudio TranslationSentence Embeddings
M
Manika Agarwal
LinkedIn, Mountain View, CA, USA
C
Chinmay Naik
LinkedIn, Mountain View, CA, USA
L
Lingyu Zhang
LinkedIn, Mountain View, CA, USA
C
Chetan Bhole
LinkedIn, Mountain View, CA, USA
C
Chirag Bhanuprasad Mehta
LinkedIn, Mountain View, CA, USA
Meng Zheng
Meng Zheng
UII America, Inc.
computer visionXAIperson re-identificationhuman modeling
P
Puneet Singh Ahluwalia
LinkedIn, Mountain View, CA, USA
S
Shirisha Singh
LinkedIn, Mountain View, CA, USA
P
Ping Jin
LinkedIn, Mountain View, CA, USA
M
Manas Apte
LinkedIn, Mountain View, CA, USA
G
Gokulraj Mohanasundaram
LinkedIn, Mountain View, CA, USA
T
Tugrul Bingol
LinkedIn, Mountain View, CA, USA
R
Raghavan Muthuregunathan
LinkedIn, Mountain View, CA, USA
Fedor Borisyuk
Fedor Borisyuk
LinkedIn
Machine learning