Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation

πŸ“… 2026-08-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing methods struggle to generate human-object interactions involving multiple objects or complex articulated structures. This work proposes a three-stage generative framework: first, it unifies the motion representation of both rigid and articulated objects using surface keypoint trajectories; second, it constructs a spatiotemporal contact distance field to model dynamic contacts between the human body and multiple objects; and third, it synthesizes full-body motions guided by these contact constraints. The approach is the first to employ surface keypoint trajectories for general object motion generation and introduces a contact representation tailored for full-body, multi-object, and articulated interaction scenarios. Evaluated on the ParaHome, HIMO, ARCTIC, and OMOMO datasets, the method achieves state-of-the-art or comparable performance across single-object, multi-object, and articulated interaction tasks.
πŸ“ Abstract
Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.
Problem

Research questions and friction points this paper is trying to address.

human-object interaction
multi-object interaction
articulated objects
motion generation
surface keypoints
Innovation

Methods, ideas, or system contributions that make the work stand out.

surface keypoint trajectories
spatio-temporal contact distance field
articulated human-object interaction
multi-object coordination
contact-guided motion synthesis