Institution profile

Spotify

Industry researcheurope · se
Official website
Research library59linked papers
Opportunities0open roles
Selected work

Representative Papers

Bayesian Inference Procedures for A/B Testing: An Overview

Aug 13, 2026

This study addresses the lack of a systematic taxonomy in Bayesian A/B testing, which has led to the conflation of prior selection and stopping rules, resulting in methodological misuse and performance risks. The authors propose a three-tier classification framework encompassing posterior consistency, error rate control under Bayes factor–based stopping, and empirical Bayes–driven false discovery rate calibration. This work provides the first comprehensive formalization of Bayesian A/B testing methodologies, demonstrating that Bayes factor stopping is approximately optimal across a range of loss functions and establishing empirical Bayes as the sole viable route to achieving third-tier calibration. Simulations reveal that flat priors combined with posterior-based stopping amount to unprincipled peeking, that well-calibrated empirical Bayes priors substantially reduce estimation error, and that expected loss–based stopping minimizes regret only when the deployment cost of null effects is negligible.

0 citationsRead paper

Hypothesis-Driven Shelf Generation for Personalised Recommendation

Jul 28, 2026

This work addresses the limitations of traditional recommendation systems, which rely on handcrafted templates and struggle to capture users’ long-tail interests. The authors propose a natural language hypothesis-driven framework for personalized shelf generation that decouples shelf planning from content retrieval. The approach comprises four stages: hypothesis generation, catalog satisfaction, shelf alignment, and offline evaluation. For the first time, large language models (LLMs) are leveraged to generate semantic hypotheses, integrated with generative retrieval, candidate selection, and an LLM-as-a-judge evaluation mechanism, enabling independent optimization of planning and retrieval. This method substantially expands the scope of personalized content provisioning and achieves user engagement on par with strong baselines in certain scenarios.

0 citationsRead paper
Recent publications

Latest Papers

Bayesian Inference Procedures for A/B Testing: An Overview

Aug 13, 2026

This study addresses the lack of a systematic taxonomy in Bayesian A/B testing, which has led to the conflation of prior selection and stopping rules, resulting in methodological misuse and performance risks. The authors propose a three-tier classification framework encompassing posterior consistency, error rate control under Bayes factor–based stopping, and empirical Bayes–driven false discovery rate calibration. This work provides the first comprehensive formalization of Bayesian A/B testing methodologies, demonstrating that Bayes factor stopping is approximately optimal across a range of loss functions and establishing empirical Bayes as the sole viable route to achieving third-tier calibration. Simulations reveal that flat priors combined with posterior-based stopping amount to unprincipled peeking, that well-calibrated empirical Bayes priors substantially reduce estimation error, and that expected loss–based stopping minimizes regret only when the deployment cost of null effects is negligible.

0 citationsRead paper

Hypothesis-Driven Shelf Generation for Personalised Recommendation

Jul 28, 2026

This work addresses the limitations of traditional recommendation systems, which rely on handcrafted templates and struggle to capture users’ long-tail interests. The authors propose a natural language hypothesis-driven framework for personalized shelf generation that decouples shelf planning from content retrieval. The approach comprises four stages: hypothesis generation, catalog satisfaction, shelf alignment, and offline evaluation. For the first time, large language models (LLMs) are leveraged to generate semantic hypotheses, integrated with generative retrieval, candidate selection, and an LLM-as-a-judge evaluation mechanism, enabling independent optimization of planning and retrieval. This method substantially expands the scope of personalized content provisioning and achieves user engagement on par with strong baselines in certain scenarios.

0 citationsRead paper