Multi-Person Human Motion Forecasting in Complex Scenes

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决复杂场景中多人运动预测问题,提出了一种结合物体信息和社交互动的条件扩散模型OCSD,通过历史动作、人物互动及物体线索综合预测,实验表明该方法有效提升了预测准确性。
📝 Abstract
Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorporating object information and social interactions into a unified framework remains particularly challenging. To address this, we propose Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework. OCSD uses an object-conditioning mechanism that modulates denoising at every timestep, enabling fine-grained human-object reasoning, and a social encoder that models the interactions between all humans in the scene. As a result, our model naturally handles varying group sizes, complex social interactions, and supports sampling multiple plausible futures. Extensive experiments show that OCSD achieves state-of-the-art results on the Humans in Kitchens (HiK) and HOI-M3 benchmarks. It reduces the two-second path error by 121.5 mm (31.3%) on HiK and 130.5 mm (33.2%) on HOI-M3 compared to prior work, and produces more realistic long-term forecasts.
Problem

Research questions and friction points this paper is trying to address.

Multi-Person Human Motion Forecasting
Complex Scenes
Object Information
Social Interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Object-Conditioned Social Diffusion
multi-person interactions
object cues
social encoder
fine-grained human-object reasoning
🔎 Similar Papers
No similar papers found.
S
Serdar Ozsoy
University of Bonn, Germany; Lamarr Institute for Machine Learning and Artificial Intelligence, Germany
L
Lars Doorenbos
University of Bonn, Germany; Lamarr Institute for Machine Learning and Artificial Intelligence, Germany
Juergen Gall
Juergen Gall
University of Bonn, Lamarr - Institute for Machine Learning and Artificial Intelligence
Action recognitionVideo understandingAnticipationForecastingHuman pose estimation