Improving Item Discoverability in e-Commerce Search via Related Intent Generation

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional e-commerce search systems, which overly rely on exact matching and consequently suffer from insufficient recall of substitute, complementary, and thematically related items, thereby hindering user discovery and commercial conversion. To overcome this, the authors propose an intent-conditioned recall expansion mechanism comprising a two-stage hybrid architecture: first, a closed-source large language model (LLM) enhances discoverability for head queries; then, a small language model (SLM), fine-tuned via LoRA and trained through teacher–student distillation, generalizes this capability to long-tail queries. This approach maintains high relevance while increasing the coverage of discoverable queries from 60% to 80% and reducing inference costs to approximately 30% of the teacher model’s, significantly boosting exposure for long-tail and emerging products and demonstrating strong effectiveness and scalability in real-world deployment.
📝 Abstract
Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of substitute, complementary, and thematically related items. In this paper, we present a scalable system for discovery-augmented search that leverages intent-conditioned recall expansion. Our approach generates implicit user intents to expand candidate recall while maintaining relevance. The system addresses the cost-quality tradeoff of generative retrieval through a two-stage hybrid architecture. First, we leverage closed-weight large language models (LLMs) to maximize discoverability for head queries. To extend these benefits to tail queries, we then introduce a finetuned small language model (SLM), trained via LoRA adapters and teacher-student distillation. We evaluate the system using a rigorous dual framework: (a) LLM-as-a-judge metrics validated against human preferences for semantic quality, and (b) end-to-end session-level purchase analysis. Results demonstrate that our approach improves both intent generation quality and downstream retrieval effectiveness, extending discovery coverage from approximately 60% to 80% of query traffic at roughly 30% of the teacher model's inference cost, offering a viable path for deployment in large-scale marketplaces. Beyond relevance gains, discovery-augmented search may serve as a marketplace-balancing mechanism, giving long-tail and emerging supply an opportunity for query-conditioned exposure.
Problem

Research questions and friction points this paper is trying to address.

item discoverability
e-commerce search
related intent
recall expansion
query matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

intent-conditioned recall expansion
discovery-augmented search
teacher-student distillation
LoRA fine-tuning
e-commerce item discoverability
🔎 Similar Papers
No similar papers found.
J
Ji Xin
Instacart
X
Xiao Xiao
Instacart
I
Ishan Bhatt
Instacart
V
Vinesh Gudla
San Francisco, USA
T
Trace Levinson
Instacart
R
Raochuan Fan
Instacart
S
Shishir Kumar Prasad
Instacart
P
Prakash Putta
Instacart
T
Tejaswi Tenneti
San Francisco, USA