🤖 AI Summary
To address the challenge of personalized, contactless fitness monitoring in home/office environments—where professional guidance and wearable devices are unavailable—this paper proposes the first smart-speaker-based acoustic sensing system for non-contact exercise monitoring. Methodologically, it integrates Doppler shift modeling, short-term energy–driven motion segmentation, and an end-to-end deep neural network, introducing a novel unified framework that jointly performs exercise action classification and user identification, with built-in incremental learning to dynamically incorporate new actions. It further defines a four-dimensional quality assessment metric encompassing duration, intensity, continuity, and fluency. Evaluated on over 9,000 repetitions of 10 exercise actions performed by 12 volunteers, the system achieves 96.13% action classification accuracy and 91% user identification accuracy, significantly enhancing autonomous training efficacy.
📝 Abstract
Fitness can help to strengthen muscles, increase resistance to diseases, and improve body shape. Nowadays, a great number of people choose to exercise at home/office rather than at the gym due to lack of time. However, it is difficult for them to get good fitness effects without professional guidance. Motivated by this, we propose the first personalized fitness monitoring system, HearFit<inline-formula><tex-math notation="LaTeX">$^+$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mo>+</mml:mo></mml:msup></mml:math><inline-graphic xlink:href="xie-ieq2-3125684.gif"/></alternatives></inline-formula>, using smart speakers at home/office. We explore the feasibility of using acoustic sensing to monitor fitness. We design a fitness detection method based on Doppler shift and adopt the short time energy to segment fitness actions. Based on deep learning, HearFit<inline-formula><tex-math notation="LaTeX">$^+$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mo>+</mml:mo></mml:msup></mml:math><inline-graphic xlink:href="xie-ieq3-3125684.gif"/></alternatives></inline-formula> can perform fitness classification and user identification at the same time. Combined with incremental learning, users can easily add new actions. We design 4 evaluation metrics (i.e., duration, intensity, continuity, and smoothness) to help users to improve fitness effects. Through extensive experiments including over 9,000 actions of 10 types of fitness from 12 volunteers, HearFit<inline-formula><tex-math notation="LaTeX">$^+$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mo>+</mml:mo></mml:msup></mml:math><inline-graphic xlink:href="xie-ieq4-3125684.gif"/></alternatives></inline-formula> can achieve an average accuracy of 96.13% on fitness classification and 91% accuracy for user identification. All volunteers confirm that HearFit<inline-formula><tex-math notation="LaTeX">$^+$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mo>+</mml:mo></mml:msup></mml:math><inline-graphic xlink:href="xie-ieq5-3125684.gif"/></alternatives></inline-formula> can help improve the fitness effect in various environments.