Institution profile

Meshcapade

Industry researcheurope · de
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

PICO: Reconstructing 3D People In Contact with Objects

Apr 24, 2025

Reconstructing 3D human–object interaction (HOI) from a single color image is challenged by depth ambiguity, severe occlusion, and high variability in object shape and appearance. Existing methods rely on controlled environments and restricted object categories, limiting generalizability. This paper introduces PICO-fit: a novel framework for open-vocabulary, end-to-end 3D HOI reconstruction from natural images. We first construct PICO-db—the first densely annotated 3D contact dataset for real-world images. Then, we propose a contact-guided render-and-compare fitting paradigm that integrates vision foundation model–based 3D object mesh retrieval, two-click contact projection, SMPL-X human body modeling, and differentiable rendering optimization. Our method achieves state-of-the-art accuracy on unseen object categories and enables the first end-to-end 3D HOI reconstruction across dozens of everyday objects. Both code and the PICO-db dataset are publicly released.

1 citationsRead paper

Open-Vocabulary Functional 3D Human-Scene Interaction Generation

Jan 28, 2026

Generating semantically and functionally plausible 3D human-scene interactions remains challenging, as existing methods often lack explicit modeling of object functionality and human contact relationships, leading to unrealistic or functionally incorrect interactions. This work proposes FunHSI, a novel framework that, for the first time, enables open-vocabulary-driven, training-free generation of functional 3D human-scene interactions. FunHSI integrates vision-language models for task understanding with functional-aware contact reasoning, 3D geometric reconstruction, and contact map modeling, followed by a staged optimization process to ensure both physical plausibility and functional correctness. Experiments demonstrate that FunHSI consistently produces realistic interactions aligned with fine-grained functional instructions—such as “sit on the sofa” or “increase room temperature”—across diverse indoor and outdoor scenes, significantly outperforming current state-of-the-art approaches.

0 citationsRead paper

FUSION: Full-Body Unified Motion Prior for Body and Hands via Diffusion

Jan 07, 2026arXiv.org

Existing full-body motion synthesis methods often neglect hand movements or generate motions only under constrained scenarios, lacking large-scale, diverse datasets that jointly capture both body and fine-grained hand articulation. To address this gap, this work integrates multi-source hand and body motion data into unified full-body motion sequences and proposes FUSION, the first unconditional diffusion-based motion prior for full-body (including fingers) synthesis. FUSION enables fine-grained interactive motions driven by either object trajectories or natural language constraints generated by large language models. Experiments demonstrate that FUSION outperforms state-of-the-art skeleton-aware control models on the HumanML3D keypoint tracking task, producing more natural motions while achieving high-precision hand control and coherent full-body coordination in both object interaction and self-interaction tasks.

0 citationsRead paper

BEDLAM2.0: Synthetic Humans and Cameras in Motion

Nov 18, 2025

This work addresses the long-standing challenge of estimating 3D human motion in world coordinates from monocular video—particularly hindered by low accuracy under synchronized human-camera motion and the scarcity of real-world ground-truth annotations. To this end, we introduce BEDLAM 2.0: the first large-scale, high-fidelity synthetic dataset for this task. Its key innovations include diverse human morphologies, clothing, hairstyles, footwear, and complex 3D environments, coupled with photorealistic camera motion trajectories. BEDLAM 2.0 provides pixel-accurate rendered videos, precise SMPL-X body parameters, and ground-truth camera poses. Extensive experiments demonstrate that models trained on BEDLAM 2.0 achieve significantly improved accuracy in world-coordinate 3D pose and motion estimation—reducing average error by 18.7% over the original BEDLAM. This establishes a robust data foundation and performance benchmark for markerless reconstruction in unconstrained real-world scenarios.

0 citationsRead paper

DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models

May 09, 2025

Single-image 3D hair reconstruction faces challenges including high hairstyle diversity, scarcity of real-world paired data, reliance on low-dimensional intermediate representations, and post-processing—limiting modeling of complex curly hairstyles (e.g., Afro). This paper proposes the first end-to-end, strand-level generative framework based on a diffusion Transformer, eliminating guided hair bundle initialization and post-hoc refinement. We introduce the largest synthetic 3D hairstyle dataset to date (40K samples), coupled with a scalp-texture-mapped latent representation and a pre-trained vision backbone to enable zero-shot generalization. Our method directly synthesizes high-fidelity, geometrically accurate individual hair strands from a single frontal image. On real images, it significantly improves curliness, density, and structural integrity over prior work. Crucially, it achieves full strand-level reconstruction without upsampling or explicit decoding—marking the first such result in the literature.

0 citationsRead paper
Recent publications

Latest Papers

Open-Vocabulary Functional 3D Human-Scene Interaction Generation

Jan 28, 2026

Generating semantically and functionally plausible 3D human-scene interactions remains challenging, as existing methods often lack explicit modeling of object functionality and human contact relationships, leading to unrealistic or functionally incorrect interactions. This work proposes FunHSI, a novel framework that, for the first time, enables open-vocabulary-driven, training-free generation of functional 3D human-scene interactions. FunHSI integrates vision-language models for task understanding with functional-aware contact reasoning, 3D geometric reconstruction, and contact map modeling, followed by a staged optimization process to ensure both physical plausibility and functional correctness. Experiments demonstrate that FunHSI consistently produces realistic interactions aligned with fine-grained functional instructions—such as “sit on the sofa” or “increase room temperature”—across diverse indoor and outdoor scenes, significantly outperforming current state-of-the-art approaches.

0 citationsRead paper

FUSION: Full-Body Unified Motion Prior for Body and Hands via Diffusion

Jan 07, 2026arXiv.org

Existing full-body motion synthesis methods often neglect hand movements or generate motions only under constrained scenarios, lacking large-scale, diverse datasets that jointly capture both body and fine-grained hand articulation. To address this gap, this work integrates multi-source hand and body motion data into unified full-body motion sequences and proposes FUSION, the first unconditional diffusion-based motion prior for full-body (including fingers) synthesis. FUSION enables fine-grained interactive motions driven by either object trajectories or natural language constraints generated by large language models. Experiments demonstrate that FUSION outperforms state-of-the-art skeleton-aware control models on the HumanML3D keypoint tracking task, producing more natural motions while achieving high-precision hand control and coherent full-body coordination in both object interaction and self-interaction tasks.

0 citationsRead paper

BEDLAM2.0: Synthetic Humans and Cameras in Motion

Nov 18, 2025

This work addresses the long-standing challenge of estimating 3D human motion in world coordinates from monocular video—particularly hindered by low accuracy under synchronized human-camera motion and the scarcity of real-world ground-truth annotations. To this end, we introduce BEDLAM 2.0: the first large-scale, high-fidelity synthetic dataset for this task. Its key innovations include diverse human morphologies, clothing, hairstyles, footwear, and complex 3D environments, coupled with photorealistic camera motion trajectories. BEDLAM 2.0 provides pixel-accurate rendered videos, precise SMPL-X body parameters, and ground-truth camera poses. Extensive experiments demonstrate that models trained on BEDLAM 2.0 achieve significantly improved accuracy in world-coordinate 3D pose and motion estimation—reducing average error by 18.7% over the original BEDLAM. This establishes a robust data foundation and performance benchmark for markerless reconstruction in unconstrained real-world scenarios.

0 citationsRead paper

DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models

May 09, 2025

Single-image 3D hair reconstruction faces challenges including high hairstyle diversity, scarcity of real-world paired data, reliance on low-dimensional intermediate representations, and post-processing—limiting modeling of complex curly hairstyles (e.g., Afro). This paper proposes the first end-to-end, strand-level generative framework based on a diffusion Transformer, eliminating guided hair bundle initialization and post-hoc refinement. We introduce the largest synthetic 3D hairstyle dataset to date (40K samples), coupled with a scalp-texture-mapped latent representation and a pre-trained vision backbone to enable zero-shot generalization. Our method directly synthesizes high-fidelity, geometrically accurate individual hair strands from a single frontal image. On real images, it significantly improves curliness, density, and structural integrity over prior work. Crucially, it achieves full strand-level reconstruction without upsampling or explicit decoding—marking the first such result in the literature.

0 citationsRead paper

PICO: Reconstructing 3D People In Contact with Objects

Apr 24, 2025

Reconstructing 3D human–object interaction (HOI) from a single color image is challenged by depth ambiguity, severe occlusion, and high variability in object shape and appearance. Existing methods rely on controlled environments and restricted object categories, limiting generalizability. This paper introduces PICO-fit: a novel framework for open-vocabulary, end-to-end 3D HOI reconstruction from natural images. We first construct PICO-db—the first densely annotated 3D contact dataset for real-world images. Then, we propose a contact-guided render-and-compare fitting paradigm that integrates vision foundation model–based 3D object mesh retrieval, two-click contact projection, SMPL-X human body modeling, and differentiable rendering optimization. Our method achieves state-of-the-art accuracy on unseen object categories and enables the first end-to-end 3D HOI reconstruction across dozens of everyday objects. Both code and the PICO-db dataset are publicly released.

1 citationsRead paper