Incremental Recommendation via Causal Models
通过扩展因果模型架构解决推荐系统中无增值推荐问题,采用双阈值策略减少7%的推荐展示,不影响内容消费。
通过扩展因果模型架构解决推荐系统中无增值推荐问题,采用双阈值策略减少7%的推荐展示,不影响内容消费。
研究探讨了通过增强SID预测和链式思维推理来改善生成推荐的自然语言描述性推理痕迹,但发现这并未提高传统离线推荐的有效性。
研究使用两种简单方法SeqRules和PCTM探究序列推荐基准是否需要高阶序列建模,发现这些基准不擅长衡量高阶序列建模带来的改进。
This study addresses the lack of a systematic taxonomy in Bayesian A/B testing, which has led to the conflation of prior selection and stopping rules, resulting in methodological misuse and performance risks. The authors propose a three-tier classification framework encompassing posterior consistency, error rate control under Bayes factor–based stopping, and empirical Bayes–driven false discovery rate calibration. This work provides the first comprehensive formalization of Bayesian A/B testing methodologies, demonstrating that Bayes factor stopping is approximately optimal across a range of loss functions and establishing empirical Bayes as the sole viable route to achieving third-tier calibration. Simulations reveal that flat priors combined with posterior-based stopping amount to unprincipled peeking, that well-calibrated empirical Bayes priors substantially reduce estimation error, and that expected loss–based stopping minimizes regret only when the deployment cost of null effects is negligible.
This work addresses the limitations of traditional recommendation systems, which rely on handcrafted templates and struggle to capture users’ long-tail interests. The authors propose a natural language hypothesis-driven framework for personalized shelf generation that decouples shelf planning from content retrieval. The approach comprises four stages: hypothesis generation, catalog satisfaction, shelf alignment, and offline evaluation. For the first time, large language models (LLMs) are leveraged to generate semantic hypotheses, integrated with generative retrieval, candidate selection, and an LLM-as-a-judge evaluation mechanism, enabling independent optimization of planning and retrieval. This method substantially expands the scope of personalized content provisioning and achieves user engagement on par with strong baselines in certain scenarios.
通过扩展因果模型架构解决推荐系统中无增值推荐问题,采用双阈值策略减少7%的推荐展示,不影响内容消费。
研究探讨了通过增强SID预测和链式思维推理来改善生成推荐的自然语言描述性推理痕迹,但发现这并未提高传统离线推荐的有效性。
研究使用两种简单方法SeqRules和PCTM探究序列推荐基准是否需要高阶序列建模,发现这些基准不擅长衡量高阶序列建模带来的改进。
This study addresses the lack of a systematic taxonomy in Bayesian A/B testing, which has led to the conflation of prior selection and stopping rules, resulting in methodological misuse and performance risks. The authors propose a three-tier classification framework encompassing posterior consistency, error rate control under Bayes factor–based stopping, and empirical Bayes–driven false discovery rate calibration. This work provides the first comprehensive formalization of Bayesian A/B testing methodologies, demonstrating that Bayes factor stopping is approximately optimal across a range of loss functions and establishing empirical Bayes as the sole viable route to achieving third-tier calibration. Simulations reveal that flat priors combined with posterior-based stopping amount to unprincipled peeking, that well-calibrated empirical Bayes priors substantially reduce estimation error, and that expected loss–based stopping minimizes regret only when the deployment cost of null effects is negligible.
This work addresses the limitations of traditional recommendation systems, which rely on handcrafted templates and struggle to capture users’ long-tail interests. The authors propose a natural language hypothesis-driven framework for personalized shelf generation that decouples shelf planning from content retrieval. The approach comprises four stages: hypothesis generation, catalog satisfaction, shelf alignment, and offline evaluation. For the first time, large language models (LLMs) are leveraged to generate semantic hypotheses, integrated with generative retrieval, candidate selection, and an LLM-as-a-judge evaluation mechanism, enabling independent optimization of planning and retrieval. This method substantially expands the scope of personalized content provisioning and achieves user engagement on par with strong baselines in certain scenarios.