🤖 AI Summary
This work addresses the challenge of capturing the inherent uncertainty in human motion within football scenarios through 3D skeletal representations. To this end, the authors propose a self-supervised representation learning framework that leverages future motion prediction as a proxy task. The approach innovatively introduces a conditional module in 3D Euclidean space to model the multimodal probability distribution of discretized future motions, thereby explicitly capturing multiple plausible motion trajectories. Experimental results demonstrate that the learned representations significantly improve prediction accuracy on large-scale football tracking data and exhibit strong cross-task generalization capabilities across diverse downstream tasks.
📝 Abstract
This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion prediction as the learning objective. Since human motion is inherently uncertain, accounting for multiple plausible futures is essential for capturing the underlying motion dynamics and learning effective representations. To this end, we introduce a conditioning module for motion prediction that models a probabilistic distribution over discretized future motions in 3D Euclidean space, learning multimodality with explicit supervision from future trajectories. Experiments on large-scale soccer player tracking data show that our approach substantially improves motion prediction accuracy. Moreover, the learned representations effectively transfer to multiple soccer downstream applications, demonstrating strong cross-task generalization.