Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过离线模型基础强化学习解决激励广告中奖励分配问题,以平衡成本与收益,并提高用户参与度和平台净利润。
📝 Abstract
Complete your ad view and grab a 5-cent bonus! In incentivized advertising, a platform promises users a bonus before observing downstream ad revenue, encouraging them to click and complete ads. It must balance the incentive promised in advance against the revenue realized afterward: insufficient incentives forfeit monetization opportunities, whereas excessive incentives reduce net profit. Because current incentives may also shape user expectations and future engagement, incentive allocation is a sequential decision problem with delayed revenue, cost sensitivity, and carryover effects. Existing work has not studied decision-making algorithms for this setting. Auto-bidding assumes available ad opportunities, while targeted promotion optimizes incentives outside the ad monetization pipeline. We formulate the problem as an MDP and develop an offline model-based RL framework for cost-controllable sequential incentive allocation. It learns a world model of user feedback and ad revenue, then performs conservative policy optimization. An independent counterfactual scorer evaluates each learned policy on held-out logs, enabling pre-launch selection without costly online exposure. Experiments on large-scale industrial data and online A/B tests show that the scorer provides a stable offline signal. The deployment path from causal inference to offline RL and then Offline-MBRL further validates the framework: MB-IQL improves per-user net profit by 7.96\% over TD3+BC, whereas reverting to plain IQL reduces it by 6.56\% (both \(p<0.0001\)).
Problem

Research questions and friction points this paper is trying to address.

incentivized advertising
sequential decision problem
delayed revenue
cost sensitivity
carryover effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

offline model-based reinforcement learning
sequential incentive allocation
conservative policy optimization
counterfactual scorer
🔎 Similar Papers
No similar papers found.
Z
Zilin Zhao
State Key Laboratory of Novel Software Technology, Nanjing University; ByteDance
H
Han Yang
State Key Laboratory of Novel Software Technology, Nanjing University; ByteDance
Tianpei Yang
Tianpei Yang
Nanjing University
Reinforcement LearningTransfer LearningMultiagent SystemsAI agents
F
Fangsheng Huang
ByteDance
Y
Yanfei Cui
ByteDance
K
Kan Peng
ByteDance
Y
Yi Li
ByteDance
Y
Yiming Zong
The Hong Kong University of Science and Technology; ByteDance
H
Hao Zhang
ByteDance
Y
Yinsong Xue
ByteDance