DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决广告推荐中技能优化的问题,本文提出文档中介强化学习(DMRL),通过结构化编辑动作序列和两个关键组件DRPO与LRP来优化技能文档,实现更有效的奖励归因与长期结果预测。
📝 Abstract
Advertising recommendation requires continuously tuning complex system parameters while balancing commercial returns and user experience. Recent work has introduced large language models (LLMs) with skill documents to assist this labor-intensive process, but skill optimization remains largely prompt-driven, lacking a principled mechanism to attribute rewards to specific document edits. To address this limitation, we propose Document-Mediated Reinforcement Learning (DMRL), a skill self-evolution framework that models skill document optimization as a sequence of structured editing actions. In DMRL, an upper-level agent performs controlled document edits, while a frozen lower-level task agent evaluates their effects through A/B testing. To address credit assignment and long-term outcomes, we introduce two key components: (1) Dual-Relative Policy Optimization (DRPO), a post-training policy optimization method for robust and risk-aware advantage estimation; and (2) Long-term Reward Predictor (LRP), which estimates long-term outcomes by modeling population heterogeneity with disentangled representation learning and cross-attention transfer. DMRL was deployed on a large-scale short-video ads platform and extensive empirical evaluation shows that DMRL outperforms state-of-the-art baselines across key advertising metrics
Problem

Research questions and friction points this paper is trying to address.

Advertising Recommendation
Skill Optimization
Document Edits
Reward Attribution
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Document-Mediated Reinforcement Learning
Dual-Relative Policy Optimization
Long-term Reward Predictor
🔎 Similar Papers
No similar papers found.