Everybody Tracking Every Body

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了多人互动场景下基于第一人称视角的3D姿态估计问题,通过融合来自个体相机运动和外部观察的姿态估计,并使用扩散方法整合数据。
📝 Abstract
We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each individual wears a camera recording egocentric video and IMU data. Processing this video with VIO SLAM provides high-quality tracking of each egocentric camera through space. The first-person view from one individual provides third-person observations of other people, although these exocentric observations are sparse, intermittent, and of highly variable reliability as both cameras and subjects move. To integrate these synchronized data streams, we propose a diffusion-based approach that fuses estimates of pose based on head motion derived from egocentric camera motion with exocentric pose observations, conditioning on both observation content and reliability. Our model is trained on a mixture of single-person motion-capture data and multi-person video in order to learn rich priors for body motion trajectories and video observation reliability. Evaluation on challenging multi-person datasets suggests our fusion approach improves over motion-only and vision-only baselines in terms of both absolute and relative pose accuracy.
Problem

Research questions and friction points this paper is trying to address.

3D body pose estimation
egocentric views
multiple interacting people
VIO SLAM
diffusion-based approach
Innovation

Methods, ideas, or system contributions that make the work stand out.

diffusion-based approach
egocentric and exocentric views fusion
observation reliability
🔎 Similar Papers
No similar papers found.