🤖 AI Summary
This work addresses the challenge of model reuse across diverse devices, environments, and modalities in heterogeneous RF sensing, which is hindered by significant discrepancies in feature structure, spatial layout, and temporal scale. To overcome this, the authors propose a unified RF representation learning framework that integrates structure-aware feature encoding, set-based spatial aggregation, and hierarchical temporal modeling. By preserving a fixed spatiotemporal backbone and adapting only the feature configuration and task-specific heads, the framework accommodates a wide range of sensing tasks. It achieves strong domain robustness, task generality, and modality scalability, attaining an average cross-domain accuracy of 92.15% across Widar3.0, CSI-Bench, and XRF55 under multifactorial settings. The approach outperforms existing methods in three of four additional tasks and reduces the cross-modal performance gap from 18.85% to 12.93%.
📝 Abstract
Heterogeneous RF sensing differs substantially in feature structure, spatial layout, and temporal scale, making existing models difficult to reuse across devices, environments, and RF modalities. We propose FSTC-Encoder, which unifies heterogeneous RF representation learning through feature, spatial, and temporal correlation modeling. Structure-aware feature encoding accommodates different signal structures, set-based spatial encoding aggregates variable observations, and hierarchical temporal encoding jointly captures local variations and long-range dependencies. Across sensing tasks and modalities, FSTC-Encoder retains the same spatial--temporal backbone architecture while varying only the feature configuration and task head. Across Widar3.0, CSI-Bench, and XRF55, FSTC-Encoder achieves 92.15% mean Accuracy under multi-factor cross-domain protocols, ranks first on three of four additional sensing tasks, remains consistently strong across WiFi, millimeter-wave radar, and RFID, and reduces the cross-modality performance gap from 18.85% to 12.93% through cross-RF learning. These results demonstrate that FSTC-Encoder achieves high domain robustness, task generality, and modality extensibility.