🤖 AI Summary
This work addresses the limited cross-user and cross-environment generalization of Wi-Fi sensing in real-world scenarios by introducing neural radiance fields to Wi-Fi-based gesture and activity recognition for the first time. The proposed framework, termed Doppler Radiance Field (DoRF++), constructs sparse observations from Doppler velocities extracted from channel state information (CSI) and leverages a neural radiance field to reconstruct the underlying 3D motion sequence. This sequence is then projected onto a unit sphere to form a structured spherical representation, which is subsequently classified using a spherical Transformer. By enabling geometric modeling of Doppler signals and spherical representation learning, DoRF++ significantly outperforms existing methods on a self-collected dataset, demonstrating notably improved accuracy—particularly for challenging gestures and in cross-user settings.
📝 Abstract
Motivated by the IEEE 802.11bf effort to standardize advanced WLAN sensing, interest in Wi-Fi Channel State Information (CSI) for passive, device-free, and privacy-preserving activity and gesture recognition has grown rapidly. Recent studies have shown that Doppler velocity projections extracted from CSI, which directly reflect human-motion velocity, enable more robust human activity recognition (HAR) and stronger generalization across users and unseen conditions. Nevertheless, reliable generalization under real-world variability remains a major challenge, hindering the adoption of Wi-Fi sensing in real-world applications. To address this challenge, we introduce Doppler Radiance Fields (DoRF), bringing the concept of neural radiance fields (NeRF) from computer vision into Wi-Fi sensing. DoRF models Doppler velocity projections extracted from Wi-Fi CSI as sparse and diverse virtual-camera views of human motion. It then infers a latent 3D motion sequence whose projections along learned effective Doppler directions explain the CSI-derived Doppler observations. The recovered motion is subsequently projected onto an equiangular grid of directions on the unit sphere, producing a spherical representation of the underlying motion. Since DoRF naturally defines the Doppler representation on spheres, we further introduce DoRF++, a spherical-learning design that applies spherical Transformers for activity classification. Experiments on our collected hand-gesture dataset show that DoRF++ significantly outperforms state-of-the-art Wi-Fi-based HAR methods in cross-user generalization accuracy, especially for difficult gestures in settings with a single multi-antenna receiver access point (AP).