Institution profile

Tripadvisor

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Significance-First Splitting: Aligning Treatment Heterogeneity Detection with Honest Estimation

Jul 04, 2026

Existing methods for heterogeneous treatment effect estimation often struggle to simultaneously ensure sensitivity in detecting effect modifiers and validity in statistical inference. This work proposes a hybrid algorithm that integrates significance-driven splitting with honest estimation: it employs the t² statistic as the splitting criterion, incorporates honest sample splitting, selects the cost-complexity penalty via cross-validation, and uses the infinitesimal jackknife to estimate Monte Carlo variance. This approach is the first to align significance-based splitting with an honest estimation framework, maintaining theoretical consistency under strong interactions and providing nominal leaf-level confidence intervals for a single tree. Empirical results demonstrate approximately 90% coverage (nominal 90%) on Athey–Imbens synthetic data and Qini coefficients on par with S- and T-learners on real-world Criteo and Starbucks datasets.

0 citationsRead paper

LLMs for estimating positional bias in logged interaction data

Sep 03, 2025

Recommendation and search systems often inherit and amplify pre-existing ranking biases due to position bias—e.g., users’ tendency to click top-ranked items—leading to distorted relevance modeling. Method: This paper proposes the first large language model (LLM)-based approach for position bias estimation, enabling fine-grained modeling of row- and column-wise positional effects under complex interface layouts directly from raw interaction logs—overcoming the expressiveness limitations of traditional heuristic methods. The LLM-generated propensity scores are integrated into an inverse propensity scoring (IPS) framework for bias correction and used to train a re-ranking model. Contribution/Results: Experiments show that, while preserving NDCG@10, the weighted NDCG@10 improves by approximately 2%, significantly mitigating layout-induced bias propagation. This work establishes a novel paradigm for trustworthy ranking.

0 citationsRead paper

Prompt Smart, Pay Less: Cost-Aware APO for Real-World Applications

Jul 18, 2025

Existing automatic prompt optimization (APO) methods struggle to adapt to high-stakes, multi-class classification tasks prevalent in commercial settings. To address this, we propose APE-OPRO—a hybrid framework integrating gradient-free Automatic Prompt Engineering (APE) and Optimized Prompt Optimization (OPRO), augmented with gradient-based techniques such as ProTeGi. It is the first APO method systematically evaluated on a real-world commercial multi-class dataset (~2,500 labeled products). Through ablation studies, we uncover the implicit sensitivity of large language models (LLMs) to label formatting. Compared to baselines, APE-OPRO reduces API cost by approximately 18% while preserving classification accuracy—outperforming OPRO and state-of-the-art approaches. It achieves an optimal trade-off between performance and computational efficiency, establishing a reproducible benchmark and practical paradigm for multi-label and multimodal prompt optimization.

0 citationsRead paper
Recent publications

Latest Papers

Significance-First Splitting: Aligning Treatment Heterogeneity Detection with Honest Estimation

Jul 04, 2026

Existing methods for heterogeneous treatment effect estimation often struggle to simultaneously ensure sensitivity in detecting effect modifiers and validity in statistical inference. This work proposes a hybrid algorithm that integrates significance-driven splitting with honest estimation: it employs the t² statistic as the splitting criterion, incorporates honest sample splitting, selects the cost-complexity penalty via cross-validation, and uses the infinitesimal jackknife to estimate Monte Carlo variance. This approach is the first to align significance-based splitting with an honest estimation framework, maintaining theoretical consistency under strong interactions and providing nominal leaf-level confidence intervals for a single tree. Empirical results demonstrate approximately 90% coverage (nominal 90%) on Athey–Imbens synthetic data and Qini coefficients on par with S- and T-learners on real-world Criteo and Starbucks datasets.

0 citationsRead paper

LLMs for estimating positional bias in logged interaction data

Sep 03, 2025

Recommendation and search systems often inherit and amplify pre-existing ranking biases due to position bias—e.g., users’ tendency to click top-ranked items—leading to distorted relevance modeling. Method: This paper proposes the first large language model (LLM)-based approach for position bias estimation, enabling fine-grained modeling of row- and column-wise positional effects under complex interface layouts directly from raw interaction logs—overcoming the expressiveness limitations of traditional heuristic methods. The LLM-generated propensity scores are integrated into an inverse propensity scoring (IPS) framework for bias correction and used to train a re-ranking model. Contribution/Results: Experiments show that, while preserving NDCG@10, the weighted NDCG@10 improves by approximately 2%, significantly mitigating layout-induced bias propagation. This work establishes a novel paradigm for trustworthy ranking.

0 citationsRead paper

Prompt Smart, Pay Less: Cost-Aware APO for Real-World Applications

Jul 18, 2025

Existing automatic prompt optimization (APO) methods struggle to adapt to high-stakes, multi-class classification tasks prevalent in commercial settings. To address this, we propose APE-OPRO—a hybrid framework integrating gradient-free Automatic Prompt Engineering (APE) and Optimized Prompt Optimization (OPRO), augmented with gradient-based techniques such as ProTeGi. It is the first APO method systematically evaluated on a real-world commercial multi-class dataset (~2,500 labeled products). Through ablation studies, we uncover the implicit sensitivity of large language models (LLMs) to label formatting. Compared to baselines, APE-OPRO reduces API cost by approximately 18% while preserving classification accuracy—outperforming OPRO and state-of-the-art approaches. It achieves an optimal trade-off between performance and computational efficiency, establishing a reproducible benchmark and practical paradigm for multi-label and multimodal prompt optimization.

0 citationsRead paper