Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决静态标题无法满足用户多样化兴趣的问题,提出GESE框架,通过生成探索和选择利用两阶段,使用LLM生成多样化标题并实时选择最优标题,提高CTR和停留时间。
📝 Abstract
In industrial recommendation feeds, presenting a static headline for an item often fails to satisfy the diverse, multimodal interests of the user population, particularly suppressing the needs of long-tail audiences. While Large Language Models (LLMs) have been integrated into recommendation for content understanding or ranking, directly optimizing them to output a single best headline typically leads to mode collapse---converging to generic patterns that satisfy average tastes but miss specific latent intents. To bridge this gap, we introduce GESE (Generate to Explore, Select to Exploit), a framework operating at the system's presentation layer that decouples personalization into generative exploration and selective exploitation. First, we treat the LLM as a probabilistic explorer, utilizing Group Sequence Policy Optimization (GSPO) with a hierarchical reward mechanism to generate a candidate set that maximizes the semantic coverage of potential user interests. Subsequently, a lightweight, real-time feedback-aware selector acts as the exploiter, identifying the optimal realization from the candidate pool based on instant contextual signals. Extensive deployment on a commercial platform with over 100 million daily active users demonstrates that GESE significantly outperforms state-of-the-art baselines, achieving a 2.57% lift in CTR and 0.87% in dwell time. These results validate that decoupling diversity-oriented generation from precision-oriented selection offers a robust blueprint for aligning generative AI with dynamic user utility.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Recommendation Feeds
Personalized Recommendation
Mode Collapse
User Interests
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generate to Explore
Select to Exploit
Group Sequence Policy Optimization
Hierarchical Reward Mechanism
Real-time Feedback-aware Selector
🔎 Similar Papers
No similar papers found.
Y
Yi Chen
Baidu Inc., Beijing, China
R
Rufeng Cheng
Baidu Inc., Beijing, China
Q
Qiang Xie
Baidu Inc., Beijing, China
T
Tao Li
Baidu Inc., Beijing, China