Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种新的自上而下的方法,通过跨模态对比学习,在统一的人体感知3D空间中实现多人重建,解决了复杂环境下多视角人体重建的难题。
📝 Abstract
Multi-view human reconstruction has been extensively studied under simplified settings, yet robust and efficient multi-person reconstruction in unconstrained environments remains challenging. Existing bottom-up methods often rely on accurate camera calibration and explicit cross-view matching, and therefore struggle with severe occlusions and ambiguities. We propose a new top-down paradigm that maintains a unified, instance-centric human-aware 3D space, enabling simultaneous camera calibration, cross-view association, and human reconstruction via cross-modal contrastive learning. Observations from multiple views are lifted and fused into this shared 3D space, where geometric structure, visual appearance, and human-centric semantic cues are jointly encoded at the instance level. We further introduce a spatial contrastive learning strategy that aligns 3D features corresponding to the same human instance across different views and modalities while separating different instances. This enables correspondence reasoning, semantic aggregation, and instance discrimination to be performed natively in 3D, improving cross-view consistency and robustness under severe occlusions. Finally, structured human body models are recovered in a feed-forward manner by regressing SMPL parameters from instance-level 3D human tokens. Extensive experiments demonstrate robust, accurate, and efficient multi-view human reconstruction in challenging real-world scenarios.
Problem

Research questions and friction points this paper is trying to address.

multi-person reconstruction
unconstrained environments
occlusions
ambiguities
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal contrastive learning
human-aware 3D space
spatial contrastive learning strategy
instance-level 3D human tokens
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuanwang Yang
College of Intelligence and Computing, Tianjin University, 135 Yaguan Road, Tianjin, 300350, Tianjin, China
Buzhen Huang
Buzhen Huang
Southeast University
Computer VisionComputer Graphics
Z
Zongxuan Ren
College of Intelligence and Computing, Tianjin University, 135 Yaguan Road, Tianjin, 300350, Tianjin, China
J
Jing Huang
College of Intelligence and Computing, Tianjin University, 135 Yaguan Road, Tianjin, 300350, Tianjin, China
Kun Li
Kun Li
Professor in Tianjin University
computer visioncomputer graphicsimage and video processing