🤖 AI Summary
This study addresses the challenges of data scarcity, strong heterogeneity, and the absence of foundational encoders in integrating large language models (LLMs) with mmWave radar. We propose a minimal textualization interface to convert point clouds into natural language and introduce mmWave-QA, the first benchmark for mmWave question answering. By integrating multi-source data, establishing a unified taxonomy, and employing calibrated perception preprocessing, this work enables standardized cross-hardware evaluation. Our research bridges the gap between radar sensing and LLMs, demonstrating their zero-shot reasoning capabilities and robustness in vision-restricted scenarios. Ultimately, this work establishes a scalable foundation for future investigations into multimodal fusion involving mmWave radar and large models.
📝 Abstract
Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for perception. While modern approaches such as vision-language models (VLMs) have attempted to incorporate reasoning capabilities into visual sensing, the integration of LLMs with the millimeter-wave (mmWave) modality-despite its unique advantages under low light and occlusion-remains largely unexplored. The principal bottlenecks stem from the scarcity of radar language pairs, severe cross-dataset heterogeneity, and the absence of a foundational mmWave encoder. We address this gap through a minimal textualization interface that serializes each mmWave point cloud into concise natural language, allowing off-the-shelf LLMs to operate in a question answering (QA) setting. Building on this, we present mmWave-QA, the first benchmark for language-conditioned mmWave human perception. mmWave-QA aggregates heterogeneous public mmWave datasets and harmonizes them via calibration-aware preprocessing and global taxonomy alignment, while providing natural language QA. Spanning six scenarios and five QA tasks, the benchmark enables standardized evaluation across diverse mmWave hardware and experimental conditions, establishing a foundation for scalable research on mmWave-LLM integration. We further evaluate and analyze LLMs on our mmWave-QA, highlighting their zero-shot reasoning potential for radar perception, as well as their robustness under visual degradation.