HandMvNet: Real-Time 3D Hand Pose Estimation Using Multi-View Cross-Attention Fusion

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出HandMvNet,通过多视角交叉注意力融合机制实时估计3D手部姿态和形状,无需相机参数输入,解决单目方法的尺度-深度模糊问题。
📝 Abstract
In this work, we present HandMvNet, one of the first real-time method designed to estimate 3D hand motion and shape from multi-view camera images. Unlike previous monocular approaches, which suffer from scale-depth ambiguities, our method ensures consistent and accurate absolute hand poses and shapes. This is achieved through a multi-view attention-fusion mechanism that effectively integrates features from multiple viewpoints. In contrast to previous multi-view methods, our approach eliminates the need for camera parameters as input to learn 3D geometry. HandMvNet also achieves a substantial reduction in inference time while delivering competitive results compared to the state-of-the-art methods, making it suitable for real-time applications. Evaluated on publicly available datasets, HandMvNet qualitatively and quantitatively outperforms previous methods under identical settings. Code is available at github.com/pyxploiter/handmvnet.
Problem

Research questions and friction points this paper is trying to address.

3D hand pose estimation
multi-view
real-time
scale-depth ambiguities
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-view cross-attention fusion
real-time 3D hand pose estimation
scale-depth ambiguities resolution
🔎 Similar Papers
M
Muhammad Asad Ali
Augmented Vision Group, German Research Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany; Department of Computer Science, University of Kaiserslautern-Landau (RPTU), Kaiserslautern, Germany
N
Nadia Robertini
Augmented Vision Group, German Research Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany
Didier Stricker
Didier Stricker
Professor for Computer Science, University Kaiserslautern
augmented realitycomputer visionimage processingbody sensor networkshci