Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

πŸ“… 2025-07-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing long-form action quality assessment (AQA) methods suffer from two key limitations: unimodal approaches neglect critical auditory cues, while multimodal methods typically employ shallow feature fusion, lacking deep cross-modal collaboration and temporal dynamic modeling. To address insufficient audio-visual synergy in artistic sports videos, this paper proposes an attention-driven multimodal alignment framework. It introduces a local query encoder for fine-grained temporal alignment, a multimodal attention consistency mechanism to enhance cross-modal interaction, and a two-level scoring scheme to improve interpretability. The model is jointly optimized via attention loss and regression loss. Extensive experiments on the RG and Fis-V datasets demonstrate significant improvements over state-of-the-art methods, validating the framework’s effectiveness, robustness, and interpretability for long-sequence AQA.

Technology Category

Application Category

πŸ“ Abstract
Long-term action quality assessment (AQA) focuses on evaluating the quality of human activities in videos lasting up to several minutes. This task plays an important role in the automated evaluation of artistic sports such as rhythmic gymnastics and figure skating, where both accurate motion execution and temporal synchronization with background music are essential for performance assessment. However, existing methods predominantly fall into two categories: unimodal approaches that rely solely on visual features, which are inadequate for modeling multimodal cues like music; and multimodal approaches that typically employ simple feature-level contrastive fusion, overlooking deep cross-modal collaboration and temporal dynamics. As a result, they struggle to capture complex interactions between modalities and fail to accurately track critical performance changes throughout extended sequences. To address these challenges, we propose the Long-term Multimodal Attention Consistency Network (LMAC-Net). LMAC-Net introduces a multimodal attention consistency mechanism to explicitly align multimodal features, enabling stable integration of visual and audio information and enhancing feature representations. Specifically, we introduce a multimodal local query encoder module to capture temporal semantics and cross-modal relations, and use a two-level score evaluation for interpretable results. In addition, attention-based and regression-based losses are applied to jointly optimize multimodal alignment and score fusion. Experiments conducted on the RG and Fis-V datasets demonstrate that LMAC-Net significantly outperforms existing methods, validating the effectiveness of our proposed approach.
Problem

Research questions and friction points this paper is trying to address.

Evaluating long-term human action quality in videos
Aligning visual and audio features for performance assessment
Capturing cross-modal interactions in extended activity sequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal attention consistency aligns visual and audio features
Local query encoder captures temporal and cross-modal relations
Attention and regression losses optimize alignment and fusion
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
X
Xin Wang
School of Sport Engineering, Beijing Sport University, Beijing, 100084, China
P
Peng-Jie Li
School of Sport Engineering, Beijing Sport University, Beijing, 100084, China
Y
Yuan-Yuan Shen
School of Sport Engineering, Beijing Sport University, Beijing, 100084, China