🤖 AI Summary
This work addresses the challenge of recovering dense 3D contact and force distributions during hand–object interactions from monocular first-person RGB images—a task that existing methods struggle to achieve. To this end, we propose EgoPHI, the first approach capable of jointly estimating dense contact maps and 3D force distributions on hand–object meshes from a single RGB image and object geometry. Our key contributions include a physics-based simulation pipeline that generates vertex-level force supervision, an end-to-end network architecture for joint contact and force estimation, and the ability to model articulated objects using only monocular input. Experiments demonstrate that EgoPHI significantly outperforms prior methods across in-distribution, out-of-distribution, and real-world scenarios, confirming its effective sim-to-real transfer capability.
📝 Abstract
Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet reasoning about physically grounded interaction requires estimating the forces acting on hands and objects, beyond localizing contact. We present EgoPHI, the first method that jointly estimates dense contact maps and 3D force distributions on hand and object meshes from a single monocular RGB image and object geometry. To address the lack of scalable ground-truth force annotations, we introduce a physics-based simulation pipeline that augments existing hand-object datasets with dense per-vertex force supervision. EgoPHI then learns dense 3D contact and force on interacting hand and articulated object meshes, extending vision-based force estimation beyond image-space or planar settings. Our evaluation on in-distribution and out-of-distribution benchmarks shows that EgoPHI improves force estimation over existing approaches while generalizing to unseen datasets. To evaluate sim-to-real transfer, we constructed two physical objects that capture dense object contact and force magnitude and used them to record a dataset of interactions from eight participants across diverse touch and grasp types. Our results demonstrate that EgoPHI recovers meaningful 3D contact and force distributions in simulated, out-of-distribution, and real-world settings, advancing egocentric hand-object understanding from contact localization toward physically grounded interaction reasoning.