🤖 AI Summary
This study addresses the challenge of accurately detecting freezing of gait (FOG) in Parkinson’s disease patients within home environments, where FOG episodes are often confounded with normal pauses and difficult to distinguish using inertial signals alone. To overcome this limitation, the work proposes a novel approach that integrates egocentric visual context from pre-trained first-person video foundation models—such as V-JEPA2—with wearable inertial measurement unit (IMU) data for enhanced FOG recognition. Evaluated via leave-one-subject-out cross-validation, the IMU-based temporal convolutional network (TCN) achieved the best performance (F1 = 42.3, AUROC = 83.0). Although egocentric vision alone did not surpass IMU-based detection, it significantly outperformed random chance, demonstrating its potential as a complementary source of contextual information in home-based clinical movement analysis.
📝 Abstract
Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly influenced by contextual factors during ADLs. Using synchronized egocentric video, wearable IMUs, and expert-annotated FOG labels collected from 13 PD participants in their homes, we evaluate frozen representations from pretrained ego-video and time-series foundation models, alongside an IMU-based TCN trained from scratch, under leave-one-subject-out evaluation. The IMU-based TCN achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC, compared with 32.6 F1 and 77.2 AUROC for V-JEPA2 ego-video features. Although ego-video alone did not outperform IMU-based sensing, it showed above-chance discrimination, and qualitative analyses suggest that egocentric vision may capture FOG-relevant information independent of IMUs. Together, these results support the use of pretrained ego-video representations to add contextual information to wearable-sensor-based clinical motion understanding in daily living.