Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了使用大型语言模型(LLMs)生成可解释的推荐理由,以提高用户对再次购买推荐的理解度,尽管这些模型在直接排名准确性上不如监督学习方法。
📝 Abstract
Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purchased items that may be needed again. In production settings, however, ranking accuracy is only one component of recommendation quality. Customers may also benefit from concise evidence about why an item is recommended now. Large language models (LLMs) offer a potential way to surface such evidence through feature-based, human-readable rationales grounded in interpretable behavioral signals. We construct repurchase features spanning cadence, frequency, recency, user behavior, and item popularity, and evaluate LLMs on two public grocery datasets and one proprietary retail dataset. We investigate (1) whether off-the-shelf LLMs can use these features as next-basket scorers relative to heuristic and supervised rankers, and (2) whether LLM-cited features carry outcome-grounded ranking signal. For the latter, we compare LLM-cited features with model-specific attribution methods under a cross-model feature-masking protocol that measures ranking degradation after masking selected features. Our results show that LLM scores are not competitive with supervised rankers, suggesting that off-the-shelf LLMs should not be used as standalone repurchase recommenders. However, changes in prompt and evidence representation can improve outcome-grounded feature-masking results in some settings even when ranking performance does not improve; the effect is dataset-dependent and does not consistently match attribution baselines. These findings suggest a practical role for LLMs as validated explanation components rather than primary rankers, with rationale quality evaluated separately from ranking accuracy.
Problem

Research questions and friction points this paper is trying to address.

Next-basket repurchase
Large language models
Feature rationales
Recommendation quality
Outcome-grounded
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Feature Rationales
Repurchase Recommendation
Outcome-Grounded Evaluation
🔎 Similar Papers
2024-05-17Annual Meeting of the Association for Computational LinguisticsCitations: 4
💼 Related Jobs
No related jobs found.
Yanan Cao
Yanan Cao
Institute of Information Engineering, Chinese Academy of Sciences
A
Anay Dombe
Walmart Global Tech, Sunnyvale, CA, United States
M
Murali Mohana Krishna Dandu
Walmart Global Tech, Sunnyvale, CA, United States
S
Shreeranjani Srirangamsridharan
Walmart Global Tech, Sunnyvale, CA, United States
S
Sinduja Subramaniam
Walmart Global Tech, Sunnyvale, CA, United States
Y
Yogananth Mahalingam
Walmart Global Tech, Sunnyvale, CA, United States
Evren Korpeoglu
Evren Korpeoglu
Walmart Global Tech
Machine learningRecommender systems
Kannan Achan
Kannan Achan
Walmartlabs
machine learningartificial intelligencegenerative modeling