Institution profile

Beijing Institute of Technology

Academic institutionasia · cn
Official website
Research library1,372linked papers
Opportunities0open roles
Selected work

Representative Papers

D3-Guard: Acoustic-based Drowsy Driving Detection Using Smartphones

Apr 01, 2019IEEE Conference on Computer Communications

This study addresses the limitation of existing fatigue-driving detection systems that rely on dedicated hardware. We propose a passive, smartphone-based acoustic sensing method leveraging built-in microphones to capture subtle Doppler-induced frequency shifts in ambient sound—caused by drowsiness-related behaviors such as yawning, head nodding, and steering wheel rotation. To enable efficient on-device processing, we introduce a lightweight undersampling–FFT feature extraction pipeline and develop an LSTM-based temporal model for early fatigue onset detection, achieving >80% detection accuracy within 70% of the behavioral event duration. To our knowledge, this is the first purely smartphone-microphone-driven acoustic fatigue detection framework. Evaluated on real-road driving data from five participants, the system achieves a mean classification accuracy of 93.31%, with low latency and strong potential for real-time, practical deployment.

42 citations4 influentialRead paper

HearSmoking: Smoking Detection in Driving Environment via Acoustic Sensing on Smartphones

Aug 01, 2022IEEE Transactions on Mobile Computing

Smoking while driving poses a significant threat to road safety, yet existing detection methods typically rely on intrusive sensors or auxiliary hardware. This paper proposes a contactless, end-to-end smoking behavior detection framework leveraging smartphone acoustic sensing: it exploits speaker–microphone co-design to emit and capture acoustic signals, capturing dynamic changes in acoustic correlation induced by the coupling of hand motion and thoracic respiration during smoking. We introduce the first temporal periodicity model characterizing composite smoking actions, enabling simultaneous hand-motion classification and respiratory rhythm analysis. The method integrates relative correlation coefficient computation, CNN-based feature learning, and explicit periodicity modeling. Evaluated under realistic driving conditions, it achieves real-time performance with an average accuracy of 93.44%. This work establishes a novel paradigm for unobtrusive, in-vehicle safety monitoring.

9 citationsRead paper

MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark

Aug 14, 2024arXiv.org

Existing mathematical reasoning benchmarks heavily rely on synthetic images, failing to capture the complexity of real-world multimodal reasoning involving photographs and textual mathematics. Method: We introduce MathScape—the first hierarchical multimodal benchmark for photorealistic mathematical problems—featuring a novel “scene–semantics–task” three-level taxonomy that systematically integrates authentic images with formal mathematical semantics, thereby addressing the longstanding gap in joint vision-language mathematical reasoning evaluation. Contribution/Results: Leveraging 11 state-of-the-art multimodal large language models (MLLMs), we conduct dual-track evaluation assessing both theoretical understanding and practical application. Empirical results reveal that even top-performing models achieve sub-50% average accuracy, exposing critical weaknesses in cross-modal alignment, symbolic parsing, and multi-step reasoning. MathScape establishes a new, high-challenge, fine-grained, and interpretable evaluation paradigm for multimodal mathematical reasoning.

7 citationsRead paper

Deep Distance Map Regression Network with Shape-Aware Loss for Imbalanced Medical Image Segmentation

Oct 04, 2020MLMI@MICCAI

To address the challenges of segmenting small targets (e.g., tumors) and severe class imbalance in medical imaging, this paper proposes a novel paradigm based on Euclidean distance map regression—replacing direct segmentation mask prediction with continuous distance map estimation followed by thresholding. Methodologically, we introduce a first-of-its-kind shape-aware joint loss function that simultaneously optimizes geometric fidelity of the predicted distance map and boundary localization accuracy. Our approach employs a U-Net variant architecture trained with a weighted L1 regression loss and a Hausdorff distance–driven boundary constraint. Evaluated on multiple benchmarks—including Skin Lesion and Cell Nuclei datasets—our method achieves consistent improvements: Dice score gains of 3.2–5.8%, a 12.7% increase in small-object recall, and a 21% reduction in boundary localization error. These results demonstrate substantial mitigation of class imbalance and enhanced capacity for geometrically accurate shape modeling.

7 citationsRead paper

HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart Speakers

May 01, 2023IEEE Transactions on Mobile Computing

To address the challenge of personalized, contactless fitness monitoring in home/office environments—where professional guidance and wearable devices are unavailable—this paper proposes the first smart-speaker-based acoustic sensing system for non-contact exercise monitoring. Methodologically, it integrates Doppler shift modeling, short-term energy–driven motion segmentation, and an end-to-end deep neural network, introducing a novel unified framework that jointly performs exercise action classification and user identification, with built-in incremental learning to dynamically incorporate new actions. It further defines a four-dimensional quality assessment metric encompassing duration, intensity, continuity, and fluency. Evaluated on over 9,000 repetitions of 10 exercise actions performed by 12 volunteers, the system achieves 96.13% action classification accuracy and 91% user identification accuracy, significantly enhancing autonomous training efficacy.

5 citationsRead paper
Recent publications

Latest Papers