DyG$^2$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DyG^2T框架,通过在粒子图上建模多尺度交互来解决从有限视觉观测中准确预测物体运动轨迹的问题。
📝 Abstract
Modeling object dynamics from limited visual observations is a fundamental problem for enabling accurate motion trajectory prediction in embodied interaction scenarios. Existing dynamics modeling methods first compress reconstructed particle representations into sparse Key Points and model their evolution using locally constrained interactions, thereby discarding fine-grained local details and obscuring discriminative interaction modeling across spatial and temporal scales, leading to drifting trajectories and inaccurate appearance prediction. To tackle these issues, we propose DyG$^2$T, a dynamics modeling framework that infers object motion trajectories by spatially completing and temporally discriminating Key Point representations and modeling multi-scale interaction over particle graphs. Spatially, DyG$^2$T enriches each Key Point by aggregating neighboring raw particle positions to recover fine-grained local details, while explicitly encoding relative offsets among Key Points to enhance geometric structure perception. Temporally, we introduce a Temporal Disentangling Network (TDN) to identify dominant cross-frame variations in latent space and amplify inter-frame differences, yielding temporally discriminative representations that are subsequently aggregated via Temporal Attention to capture frame-wise temporal evolution cues. For comprehensive interaction modeling, a Particle Graph Transformer leverages global attention to preserve discriminative long-range dependencies among Key Points, mitigating representation homogenization induced by locality-constrained modeling and providing a robust basis for accurate trajectory prediction. Experiments on both synthetic and real-world datasets demonstrate that DyG$^2$T achieves accurate dynamics modeling and reasoning, and exhibits strong cross-object and real-world generalization.
Problem

Research questions and friction points this paper is trying to address.

object dynamics
limited visual observations
motion trajectory prediction
spatial and temporal scales
trajectory drifting
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Gaussian Temporal-Spatial Particle Graph
Temporal Disentangling Network (TDN)
Particle Graph Transformer
Y
Yansong Wang
School of Computer Science and Technology, Harbin Institute of Technology at Weihai, Weihai 264209, China
Zhaobo Qi
Zhaobo Qi
HIT
video understandingmultimodal reasoning
X
Xinyan Liu
School of Computer Science and Technology, Harbin Institute of Technology at Weihai, Weihai 264209, China
Beichen Zhang
Beichen Zhang
School of Computer Science & Technology, University of Chinese Academy of Sciences
machine learningartificial intelligence
S
Shuhui Wang
Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China
Weigang Zhang
Weigang Zhang
Professor of Computer Science, Harbin Institute of Technology, Weihai
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer Vision
Qingming Huang
Qingming Huang
University of the Chinese Academy of Sciences
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer VisionVideo Coding