FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of systematic evaluation of multimodal large language models (MLLMs) in assessing exercise form quality (Action Quality Assessment, AQA), where existing benchmarks suffer from inconsistent error definitions and insufficient granularity. To bridge this gap, we introduce FitAQA, a novel benchmark comprising 2,219 videos across 30 bodyweight exercises and 5,512 question-answer pairs, grounded in a six-dimensional form error ontology collaboratively developed with sports science experts. FitAQA encompasses three core tasks—perception, judgment, and temporal localization—enabling, for the first time, fine-grained and interpretable AQA evaluation tailored to MLLMs. Experimental results reveal that current models still struggle with comprehensive assessment and precise error localization, with visual perception identified as the primary bottleneck; notably, incorporating real perceptual evidence substantially improves judgment accuracy.
📝 Abstract
Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored. Existing benchmarks rely on action-specific annotation schemes and focus primarily on final assessment outputs, offering limited insight into how models assess exercise quality. We introduce FitAQA, a systematic benchmark for evaluating MLLMs in fitness AQA, containing 2,219 videos and 5,512 QA instances across 30 bodyweight exercises. In collaboration with experts in sports science, we develop a unified form error taxonomy that defines 38 recurring form errors within six complementary quality dimensions: alignment, symmetry, stability, coordination, tempo, and completeness. This taxonomy provides a shared assessment framework across different exercises. FitAQA further formulates three evaluation tasks: perception for recognizing relevant visual evidence, judgement for combining that evidence with domain knowledge to assess execution correctness, and temporal grounding for localizing form errors over time. Extensive evaluation shows that current MLLMs still struggle to assess exercise quality comprehensively and localize form errors precisely. Controlled experiments further indicate that visual perception is a key bottleneck, as judgement performance improves substantially when ground-truth perceptual evidence is provided. The dataset and evaluation code will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

Fitness Action Quality Assessment
Multimodal Large Language Models
form error
benchmark
sports training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fitness Action Quality Assessment
Multimodal Large Language Models
Form Error Taxonomy
Temporal Grounding
Systematic Benchmark
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
K
Kaili Zheng
Department of Electronic Engineering, Tsinghua University
K
Kaiwen Wang
Department of Electronic Engineering, Tsinghua University
Xun Zhu
Xun Zhu
Tsinghua University
Multi-modal LLMMulti-task LearningSpatio-temporal Forecasting
Q
Qingyuan Yang
Xinjiang Region Sports Science Research Center
C
Chenyi Guo
Department of Electronic Engineering, Tsinghua University
Ji Wu
Ji Wu
Tsinghua University
Artificial Intelligence,smart healthcaremachine learningpattern recognitionspeech recognition