World Model-Guided Reinforcement Learning via Counterfactual User Engagement Simulation

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种通过反事实用户参与模拟的世界模型引导的强化学习方法,以解决用户中心代理在线反馈成本高、延迟大和缺乏相同用户状态下反事实比较的问题。
📝 Abstract
Reinforcement learning for user-centric agents is limited by the cost, latency, and risk of collecting online feedback, as well as by the lack of counterfactual comparisons under the same user state. In this paper, we propose World Model-Guided Reinforcement Learning via counterfactual user engagement simulation (WMG-RL), a framework in which a frozen user simulator provides reward supervision before real user exposure. Motivated by language world models, we instantiate the simulator as a User Engagement World Model (UEWM), which treats a recommended item as the agent action and the user's heterogeneous feedback as the environment observation. Rather than learning one fixed environment transition, UEWM learns to infer user-specific dynamics from engagement history and apply them to candidate items. In WMG-RL, a downstream policy proposes multiple candidate items for the same history; UEWM predicts the corresponding engagement feedback in parallel; and the simulated feedback is converted into dense rewards for policy optimization. Experiments show that UEWM provides reliable and transferable reward signals across domains, and that WMG-RL enables a compact 1.7B student policy to match or surpass much larger LLMs on downstream recommendation tasks.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
User-centric Agents
Online Feedback
Counterfactual Comparisons
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Model-Guided Reinforcement Learning
Counterfactual User Engagement Simulation
User Engagement World Model
Dense Rewards
Policy Optimization
🔎 Similar Papers
No similar papers found.
A
Ang Li
The Chinese University of Hong Kong; MoE Key Lab of High Confidence Software Technologies, CUHK
X
Xin Xu
ByteDance China
Bin Liang
Bin Liang
The Chinese University of Hong Kong
NLPdata miningmachine learningtext mining
Yue Ma
Yue Ma
Bytedance
NLPDialogue SystemLLM
F
Fubang Zhao
ByteDance China
Yangyang Kang
Yangyang Kang
DAMO Academy, Alibaba Group
LLM NLP KG DL
K
Kam-Fai Wong
The Chinese University of Hong Kong; MoE Key Lab of High Confidence Software Technologies, CUHK