MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
为解决动作质量评估中忽视肌肉力学等问题,MyoMechanix通过多模态数据和结构化表示方法提供细粒度的生物力学反馈。
为解决动作质量评估中忽视肌肉力学等问题,MyoMechanix通过多模态数据和结构化表示方法提供细粒度的生物力学反馈。
Existing long-form action quality assessment (AQA) methods suffer from two key limitations: unimodal approaches neglect critical auditory cues, while multimodal methods typically employ shallow feature fusion, lacking deep cross-modal collaboration and temporal dynamic modeling. To address insufficient audio-visual synergy in artistic sports videos, this paper proposes an attention-driven multimodal alignment framework. It introduces a local query encoder for fine-grained temporal alignment, a multimodal attention consistency mechanism to enhance cross-modal interaction, and a two-level scoring scheme to improve interpretability. The model is jointly optimized via attention loss and regression loss. Extensive experiments on the RG and Fis-V datasets demonstrate significant improvements over state-of-the-art methods, validating the framework’s effectiveness, robustness, and interpretability for long-sequence AQA.
为解决动作质量评估中忽视肌肉力学等问题,MyoMechanix通过多模态数据和结构化表示方法提供细粒度的生物力学反馈。
Existing long-form action quality assessment (AQA) methods suffer from two key limitations: unimodal approaches neglect critical auditory cues, while multimodal methods typically employ shallow feature fusion, lacking deep cross-modal collaboration and temporal dynamic modeling. To address insufficient audio-visual synergy in artistic sports videos, this paper proposes an attention-driven multimodal alignment framework. It introduces a local query encoder for fine-grained temporal alignment, a multimodal attention consistency mechanism to enhance cross-modal interaction, and a two-level scoring scheme to improve interpretability. The model is jointly optimized via attention loss and regression loss. Extensive experiments on the RG and Fis-V datasets demonstrate significant improvements over state-of-the-art methods, validating the framework’s effectiveness, robustness, and interpretability for long-sequence AQA.