Data Fusion-Enhanced Decision Transformer for Stable Cross-Domain Generalization

📅 2025-11-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Decision Transformers (DTs) suffer from semantic discontinuity and degraded reasoning in cross-domain transfer due to source trajectory stitching mismatches—causing state structural misalignment, incomparable return-to-go (RTG) values, and action jumps. To address this, we propose a data-fusion augmentation framework: (i) a two-tier filtering mechanism jointly leveraging Maximum Mean Discrepancy (MMD) and Optimal Transport (OT) to align both state structure and action feasibility; (ii) advantage-conditioned tokens replacing RTG tokens to enhance sequential semantic consistency; and (iii) Q-guided value regularization coupled with feasibility-weighted distributional training to improve generalization stability. We theoretically establish a bound linking transfer performance to distribution matching error. Evaluated on D4RL’s multi-domain benchmarks, our method significantly outperforms both offline RL and sequence modeling baselines, achieving superior cumulative returns and policy robustness under domain shifts in gravity, kinematics, and morphology.

Technology Category

Application Category

📝 Abstract
Cross-domain shifts present a significant challenge for decision transformer (DT) policies. Existing cross-domain policy adaptation methods typically rely on a single simple filtering criterion to select source trajectory fragments and stitch them together. They match either state structure or action feasibility. However, the selected fragments still have poor stitchability: state structures can misalign, the return-to-go (RTG) becomes incomparable when the reward or horizon changes, and actions may jump at trajectory junctions. As a result, RTG tokens lose continuity, which compromises DT's inference ability. To tackle these challenges, we propose Data Fusion-Enhanced Decision Transformer (DFDT), a compact pipeline that restores stitchability. Particularly, DFDT fuses scarce target data with selectively trusted source fragments via a two-level data filter, maximum mean discrepancy (MMD) mismatch for state-structure alignment, and optimal transport (OT) deviation for action feasibility. It then trains on a feasibility-weighted fusion distribution. Furthermore, DFDT replaces RTG tokens with advantage-conditioned tokens, which improves the continuity of the semantics in the token sequence. It also applies a $Q$-guided regularizer to suppress junction value and action jumps. Theoretically, we provide bounds that tie state value and policy performance gaps to the MMD-mismatch and OT-deviation measures, and show that the bounds tighten as these two measures shrink. We show that DFDT improves return and stability over strong offline RL and sequence-model baselines across gravity, kinematic, and morphology shifts on D4RL-style control tasks, and further corroborate these gains with token-stitching and sequence-semantics stability analyses.
Problem

Research questions and friction points this paper is trying to address.

Addresses cross-domain policy failures due to state-action misalignment and reward shifts
Enhances trajectory stitchability by fusing target data with filtered source fragments
Improves token sequence continuity through advantage conditioning and value regularization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fuses target data with source fragments via two-level filter
Replaces RTG tokens with advantage-conditioned tokens
Applies Q-guided regularizer to suppress junction jumps
🔎 Similar Papers
No similar papers found.
Guojian Wang
Guojian Wang
Luxium Solutions
Crystal growthSemiconductorlaseroptical materialshigh purity germanium growth and detector fabrication
Q
Quinson Hon
The Chinese University of Hong Kong
X
Xuyang Chen
National University of Singapore
L
Lin Zhao
National University of Singapore