🤖 AI Summary
This study addresses the challenges of scarce, small-scale, and poorly reproducible behavioral data in XR by proposing a motion synthesis pipeline based on Dynamic Time Warping and trajectory interpolation. Innovatively treating synthetic data as an independent behavioral population rather than mere augmentation, this approach generates large-scale synthetic trajectories that preserve task structure while remaining distinguishable. Experiments demonstrate that the released 100 synthetic trajectories exhibit low confusion with real data. Furthermore, mixed datasets achieve user identification performance comparable to purely real datasets. These findings effectively mitigate data bottlenecks in XR behavioral modeling and provide high-quality data support for machine learning evaluation, establishing synthetic data as a viable alternative when real-world collection is constrained.
📝 Abstract
Large-scale behavioral datasets are becoming increasingly important for machine learning, personalization, and behavioral modeling in extended reality (XR). However, collecting XR motion data from hundreds or thousands of participants remains expensive, time-consuming, and difficult to reproduce across research groups. As a result, many XR studies continue to rely on relatively small datasets that limit the scale and diversity of behavioral evaluation. To address this limitation, we investigate synthetic behavioral populations as a complementary approach to traditional XR data collection. We present an interpolation-based motion synthesis pipeline that combines dynamic time warping (DTW) with trajectory interpolation to generate synthetic behavioral trajectories from existing XR datasets while preserving task structure and incorporating motion characteristics from contributing participants. Using the publicly available FAST VR assembly dataset, we generated and openly released 100 synthetic behavioral trajectories. We evaluated the synthesized trajectories through motion-based user identification. Hybrid datasets containing both real and synthesized trajectories achieved performance comparable to similarly sized real-only datasets while maintaining low confusion between synthesized trajectories and their contributing participants. Rather than serving as conventional data augmentation, the proposed approach generates distinguishable behavioral trajectories that expand XR behavioral populations for larger-scale behavioral modeling and machine learning evaluation. Our findings demonstrate that synthetic behavioral populations provide a promising approach to expanding XR behavioral datasets and supporting future data-driven immersive systems.