The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce
本文提出差分推理路由器解决电商中冷启动LLM标注问题,通过成本感知框架优化模型选择与人工介入,提高标注效率和准确性。
本文提出差分推理路由器解决电商中冷启动LLM标注问题,通过成本感知框架优化模型选择与人工介入,提高标注效率和准确性。
Existing e-commerce visual search systems suffer from limited generalization and scalability due to category-coupled detection-classification pipelines and reliance on noisy labels. This work proposes a category-decoupled visual search architecture that employs category-agnostic region proposals and a unified embedding space for similarity retrieval. To eliminate dependence on manual annotations or noisy catalog data, we introduce a large language model–based zero-shot evaluation mechanism (LLM-as-a-Judge). The proposed approach significantly enhances system robustness and generalization. Upon large-scale deployment on a real-world home furnishings e-commerce platform, it yields substantial improvements in retrieval quality and user engagement, with offline evaluation metrics showing strong alignment with online performance.
E-commerce review summarization suffers from insufficient interpretability and practical utility. Method: This paper proposes a guided summarization framework integrating Aspect-Based Sentiment Analysis (ABSA) with Large Language Models (LLMs). It first extracts high-frequency aspect–sentiment pairs and samples representative reviews to construct structured prompts, enabling LLMs to generate faithful, concise, and attributable summaries. Second, it introduces a lightweight real-time inference architecture for scalable online deployment. Contributions/Results: (1) The first large-scale public review dataset for home-eCommerce—comprising 11.8 million anonymized reviews across 92,000 products; (2) Significant improvements in user click-through rate and satisfaction, validated via rigorous A/B testing; (3) An open-source, reproducible ABSA+LLM collaborative summarization paradigm that jointly ensures accuracy, interpretability, and engineering scalability.
In multi-treatment, partially overlapping observational causal inference settings—such as multi-channel marketing or multi-category ordering—conventional approaches that estimate treatment effects independently suffer from high variance and weak decision support. This paper proposes a customized ridge regression framework that employs a data-driven shrinkage strategy to adaptively balance heterogeneity and homogeneity across treatment effects. The method reduces the mean squared error of individual effect estimates while preserving interpretability and reconstructibility of aggregated causal effects. Theoretical analysis establishes consistency and convergence of the estimator; simulation studies corroborate its finite-sample performance. Real-world deployment at Wayfair demonstrates substantial improvements in causal decision accuracy and operational reliability for multi-treatment scenarios.
本文提出差分推理路由器解决电商中冷启动LLM标注问题,通过成本感知框架优化模型选择与人工介入,提高标注效率和准确性。
Existing e-commerce visual search systems suffer from limited generalization and scalability due to category-coupled detection-classification pipelines and reliance on noisy labels. This work proposes a category-decoupled visual search architecture that employs category-agnostic region proposals and a unified embedding space for similarity retrieval. To eliminate dependence on manual annotations or noisy catalog data, we introduce a large language model–based zero-shot evaluation mechanism (LLM-as-a-Judge). The proposed approach significantly enhances system robustness and generalization. Upon large-scale deployment on a real-world home furnishings e-commerce platform, it yields substantial improvements in retrieval quality and user engagement, with offline evaluation metrics showing strong alignment with online performance.
E-commerce review summarization suffers from insufficient interpretability and practical utility. Method: This paper proposes a guided summarization framework integrating Aspect-Based Sentiment Analysis (ABSA) with Large Language Models (LLMs). It first extracts high-frequency aspect–sentiment pairs and samples representative reviews to construct structured prompts, enabling LLMs to generate faithful, concise, and attributable summaries. Second, it introduces a lightweight real-time inference architecture for scalable online deployment. Contributions/Results: (1) The first large-scale public review dataset for home-eCommerce—comprising 11.8 million anonymized reviews across 92,000 products; (2) Significant improvements in user click-through rate and satisfaction, validated via rigorous A/B testing; (3) An open-source, reproducible ABSA+LLM collaborative summarization paradigm that jointly ensures accuracy, interpretability, and engineering scalability.
In multi-treatment, partially overlapping observational causal inference settings—such as multi-channel marketing or multi-category ordering—conventional approaches that estimate treatment effects independently suffer from high variance and weak decision support. This paper proposes a customized ridge regression framework that employs a data-driven shrinkage strategy to adaptively balance heterogeneity and homogeneity across treatment effects. The method reduces the mean squared error of individual effect estimates while preserving interpretability and reconstructibility of aggregated causal effects. Theoretical analysis establishes consistency and convergence of the estimator; simulation studies corroborate its finite-sample performance. Real-world deployment at Wayfair demonstrates substantial improvements in causal decision accuracy and operational reliability for multi-treatment scenarios.