Institution profile

Deemos Technology

Industry research
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos

Mar 18, 2026

This work addresses the challenge of automatically generating realistic and view-consistent appearances for untextured 3D models—a critical task in digital content creation. The authors propose a geometry-conditioned video diffusion approach that leverages multimodal geometric feature encoding to constrain the generation of 360-degree turntable videos. A novel 3D-aware inpainting mechanism is introduced to reconstruct self-occluded regions, ensuring complete surface coverage. By explicitly integrating geometric priors into the video diffusion framework, the method produces high-quality, temporally coherent dynamic previews suitable for direct use in UV texture back-projection or as supervision for neural rendering pipelines. Experimental results demonstrate significant improvements over existing techniques in both view consistency and 3D reconstruction fidelity, enabling fully automated production of ready-to-use 3D assets.

0 citationsRead paper

ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K

Mar 17, 2026

This work addresses the scarcity of large-scale, diverse, and simulation-ready 3D digital assets for robot learning by introducing ManiTwin—an end-to-end automated pipeline that generates physically plausible, semantically annotated, linguistically described, and manipulation-verifiable 3D object twins from a single input image. Integrating image-to-3D reconstruction, semantic and functional analysis, physics-based modeling, and manipulation feasibility verification, the proposed method enables the creation of ManiTwin-100K, a novel dataset comprising 100,000 high-quality assets. This dataset substantially advances the scale and efficiency of simulation training data generation, thereby supporting a range of downstream tasks including robotic manipulation policy learning, stochastic scene synthesis, and vision-language question answering.

0 citationsRead paper

Mojito: LLM-Aided Motion Instructor with Jitter-Reduced Inertial Tokens

Feb 22, 2025

Existing vision/audio-based multimodal motion understanding methods struggle to model 3D dynamic forces and torques; while IMUs offer lightweight, privacy-preserving advantages, their utility for long-term, real-time motion capture (MoCap) and online analysis is hindered by wireless transmission instability, sensor noise, and drift. This paper introduces the first LLM-driven inertial motion instruction framework. It proposes jitter-suppressed inertial token representations to enable noise-robust temporal modeling and semantic motion parsing. The framework integrates a lightweight IMU array, a jitter-aware encoder, an LLM-enhanced motion–language alignment architecture, and edge–cloud collaborative streaming inference. Evaluated in real-world settings, it achieves 92.3% action intention recognition accuracy with sub-80 ms latency, reduces drift error by 67%, and supports 12-hour calibration-free continuous MoCap with real-time behavioral feedback.

0 citationsRead paper

CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image

Feb 18, 2025

Single-image 3D scene reconstruction faces longstanding challenges including occlusion, geometric interpenetration, and physically implausible configurations. To address these, we propose an end-to-end framework: (1) joint 2D semantic segmentation and relative depth estimation for initial geometry; (2) GPT-driven spatial relation modeling to explicitly reason about object pose and support relationships; (3) occlusion-aware, part-level 3D generation via MAE-based point cloud conditioning; and (4) joint optimization using Signed Distance Fields (SDFs) fused with physics-aware constraint graphs to resolve floating, interpenetration, and support failures. We introduce the novel “component alignment” paradigm, enforcing semantic, geometric, and physical consistency simultaneously. Our method generates high-fidelity, texture-coherent, and physically plausible complete 3D scenes from real-world images, significantly improving object localization accuracy and cross-object spatial coherence—enabling downstream applications such as robotic simulation.

0 citationsRead paper

TANGLED: Generating 3D Hair Strands from Images with Arbitrary Styles and Viewpoints

Feb 10, 2025

Existing 3D hair generation methods struggle to model hairstyle diversity and geometric complexity—particularly for braided styles, cross-cultural aesthetics, and low-quality inputs. To address this, we introduce MultiHair, the first large-scale dataset explicitly designed for hairstyle diversity, and propose a sketch-guided multi-view diffusion framework built upon latent diffusion models (LDMs). Our method integrates multi-view sketch encoding, topology-aware conditional modeling, cross-attention mechanisms, and a parametric braid constraint module. It supports both single- and multi-view image inputs and synthesizes high-fidelity, view- and style-arbitrary 3D hair strand models. Extensive experiments demonstrate that our approach significantly outperforms prior methods in realism, structural consistency, and cultural expressiveness. Notably, it achieves breakthrough performance on complex braided hairstyles and exhibits superior robustness to input degradation.

0 citationsRead paper
Recent publications

Latest Papers

TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos

Mar 18, 2026

This work addresses the challenge of automatically generating realistic and view-consistent appearances for untextured 3D models—a critical task in digital content creation. The authors propose a geometry-conditioned video diffusion approach that leverages multimodal geometric feature encoding to constrain the generation of 360-degree turntable videos. A novel 3D-aware inpainting mechanism is introduced to reconstruct self-occluded regions, ensuring complete surface coverage. By explicitly integrating geometric priors into the video diffusion framework, the method produces high-quality, temporally coherent dynamic previews suitable for direct use in UV texture back-projection or as supervision for neural rendering pipelines. Experimental results demonstrate significant improvements over existing techniques in both view consistency and 3D reconstruction fidelity, enabling fully automated production of ready-to-use 3D assets.

0 citationsRead paper

ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K

Mar 17, 2026

This work addresses the scarcity of large-scale, diverse, and simulation-ready 3D digital assets for robot learning by introducing ManiTwin—an end-to-end automated pipeline that generates physically plausible, semantically annotated, linguistically described, and manipulation-verifiable 3D object twins from a single input image. Integrating image-to-3D reconstruction, semantic and functional analysis, physics-based modeling, and manipulation feasibility verification, the proposed method enables the creation of ManiTwin-100K, a novel dataset comprising 100,000 high-quality assets. This dataset substantially advances the scale and efficiency of simulation training data generation, thereby supporting a range of downstream tasks including robotic manipulation policy learning, stochastic scene synthesis, and vision-language question answering.

0 citationsRead paper

Mojito: LLM-Aided Motion Instructor with Jitter-Reduced Inertial Tokens

Feb 22, 2025

Existing vision/audio-based multimodal motion understanding methods struggle to model 3D dynamic forces and torques; while IMUs offer lightweight, privacy-preserving advantages, their utility for long-term, real-time motion capture (MoCap) and online analysis is hindered by wireless transmission instability, sensor noise, and drift. This paper introduces the first LLM-driven inertial motion instruction framework. It proposes jitter-suppressed inertial token representations to enable noise-robust temporal modeling and semantic motion parsing. The framework integrates a lightweight IMU array, a jitter-aware encoder, an LLM-enhanced motion–language alignment architecture, and edge–cloud collaborative streaming inference. Evaluated in real-world settings, it achieves 92.3% action intention recognition accuracy with sub-80 ms latency, reduces drift error by 67%, and supports 12-hour calibration-free continuous MoCap with real-time behavioral feedback.

0 citationsRead paper

CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image

Feb 18, 2025

Single-image 3D scene reconstruction faces longstanding challenges including occlusion, geometric interpenetration, and physically implausible configurations. To address these, we propose an end-to-end framework: (1) joint 2D semantic segmentation and relative depth estimation for initial geometry; (2) GPT-driven spatial relation modeling to explicitly reason about object pose and support relationships; (3) occlusion-aware, part-level 3D generation via MAE-based point cloud conditioning; and (4) joint optimization using Signed Distance Fields (SDFs) fused with physics-aware constraint graphs to resolve floating, interpenetration, and support failures. We introduce the novel “component alignment” paradigm, enforcing semantic, geometric, and physical consistency simultaneously. Our method generates high-fidelity, texture-coherent, and physically plausible complete 3D scenes from real-world images, significantly improving object localization accuracy and cross-object spatial coherence—enabling downstream applications such as robotic simulation.

0 citationsRead paper

TANGLED: Generating 3D Hair Strands from Images with Arbitrary Styles and Viewpoints

Feb 10, 2025

Existing 3D hair generation methods struggle to model hairstyle diversity and geometric complexity—particularly for braided styles, cross-cultural aesthetics, and low-quality inputs. To address this, we introduce MultiHair, the first large-scale dataset explicitly designed for hairstyle diversity, and propose a sketch-guided multi-view diffusion framework built upon latent diffusion models (LDMs). Our method integrates multi-view sketch encoding, topology-aware conditional modeling, cross-attention mechanisms, and a parametric braid constraint module. It supports both single- and multi-view image inputs and synthesizes high-fidelity, view- and style-arbitrary 3D hair strand models. Extensive experiments demonstrate that our approach significantly outperforms prior methods in realism, structural consistency, and cultural expressiveness. Notably, it achieves breakthrough performance on complex braided hairstyles and exhibits superior robustness to input degradation.

0 citationsRead paper