Learning human joint torques from pixels

๐Ÿ“… 2026-08-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work proposes the first method for estimating human joint torques from monocular RGB videos without relying on electromyography, motion capture markers, force plates, or simulation data. To this end, the authors introduce the VID dataset and benchmark, which integrates real-world images, kinematic annotations, anthropometric parameters, and OpenSim-generated dynamic labels. They also design the VID-Network, an end-to-end model that leverages pose-pretrained spatial probabilistic features, marker regression, and temporal torque reasoning to predict joint torques. This study establishes the first practical vision-driven benchmark for human inverse dynamics and defines a standardized evaluation protocol encompassing whole-body, joint-level, and action-level metrics. Experiments show that VID-Network achieves a mean normalized joint torque error (mPJE) of 1.7612 Nยทm/kg on VID, outperforming the best baseline by 39.81% and demonstrating superior performance across diverse joints and most actions.
๐Ÿ“ Abstract
Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-world movement scenarios. Existing torque estimation methods typically depend on surface electromyography, motion-capture markers, force plates, or simulated imitation data, which limits their applicability to ordinary RGB images. In this work, we introduce VID, a vision-based inverse dynamics dataset and benchmark for predicting human joint torques directly from real monocular images. VID contains 63,369 synchronized frames with real human images, kinematic annotations, anthropometric attributes, and OpenSim-derived dynamic labels, providing paired visual and biomechanical supervision for real-image inverse dynamics. We further define a standardized evaluation protocol covering overall torque estimation, joint-specific analysis, and action-specific prediction. To establish a strong reference model, we propose VID-Network, which combines pose-pretrained spatial probabilistic features, marker regression, and temporal torque inference to recover joint torques from image sequences. Experiments on VID show that VID-Network achieves an overall mPJE of 1.7612 N$\cdot$m/kg, improving over the best compared baseline by 39.81\%, and obtains the lowest error across all evaluated joint types and most action categories. VID establishes a first practical benchmark for vision-driven human inverse dynamics and provides a foundation for studying biomechanical inference in less constrained environments.
Problem

Research questions and friction points this paper is trying to address.

human joint torques
vision-based inverse dynamics
biomechanical analysis
monocular images
real-world movement
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-based inverse dynamics
joint torque estimation
monocular RGB images
biomechanical analysis
VID-Network
๐Ÿ”Ž Similar Papers
No similar papers found.