CoDrift: Compositional Drifting for Offline Reinforcement Learning

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CoDrift框架,通过组合三种目标层面的运动场为统一策略场,解决离线强化学习中保持数据集行为支持与选择高价值动作的多目标问题。
📝 Abstract
Offline reinforcement learning is intrinsically multi-objective: a policy must remain compatible with the behavioral support of a fixed dataset while preferentially selecting high-value actions. We recast these objectives in a common form by viewing each as an action-space motion field that specifies how generated actions should move. This perspective enables heterogeneous learning objectives to be combined directly through field composition. Inspired by drifting models, we propose CoDrift, a compositional framework for one-step generative policy learning. CoDrift combines three objective-level fields into a unified policy field. The conditional field preserves state-dependent behavioral structure, while the marginal field pools actions across states to provide a more stable generative signal in the single-positive-sample regime of continuous-control offline RL. The value field moves generated actions toward higher-value regions. The composed field is absorbed into a stochastic generator that produces an action with a single forward pass at deployment. We evaluate CoDrift on 73 tasks from OGBench and D4RL in both offline and offline-to-online settings. CoDrift compares favorably with state-of-the-art methods and achieves the best average rank in both settings.
Problem

Research questions and friction points this paper is trying to address.

Offline Reinforcement Learning
Multi-objective
Policy Learning
Action-space Motion Field
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Drifting
Offline Reinforcement Learning
Action-Space Motion Field
Generative Policy Learning
Stochastic Generator
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
X
Xiewei Ni
Xi’an Jiaotong University
R
Ruofeng Mei
Xi’an Jiaotong University
Xiangyu Xu
Xiangyu Xu
Xi’an Jiaotong University
Computer VisionMachine LearningImage Processing