Institution profile

World Labs

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

Jun 11, 2026

Existing image-to-3D methods struggle to simultaneously achieve pixel-aligned fidelity and complete scene geometry. To address this, this work proposes World Tracing—a generative, pixel-aligned geometric representation that predicts, for each input pixel, an ordered stack of multi-layer 3D points in camera space, where the first layer corresponds to the visible surface and subsequent layers reconstruct occluded structures. This approach uniquely unifies complete geometric generation with precise pixel alignment while preserving 2D–3D correspondences, enabling text-driven editing, novel view synthesis, and training-free integration with textured mesh generators. Built upon a World Tracing Diffusion Transformer (WT-DiT), the method employs factorized yet globally attentive mechanisms to denoise multi-layer geometry, trained via pixel-space flow matching and a hybrid noise schedule. Experiments across object, scene, and dynamic datasets demonstrate significant improvements over current depth estimation and image-to-3D generation models in both visible surface reconstruction and holistic geometry completion.

0 citationsRead paper

Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

Jun 11, 2026

This work addresses the challenge of efficiently generating dynamic 4D human representations from monocular or sparse multi-view videos without relying on explicit geometric priors such as skeletons or depth maps. The authors propose a diffusion-based multi-view video generation method conditioned solely on relative camera poses, integrating spatiotemporal information and SE(3) camera geometry through a five-axis positional encoding scheme. A three-stage curriculum learning strategy enables flexible novel-view synthesis and temporal extrapolation. Built upon the Wan 2.1 1.3B text-to-video diffusion model, the approach incorporates extended spatiotemporal RoPE, historical target-view tokens, and multi-view textual conditioning. Evaluated on the DNA-Rendering and ActorsHQ datasets, the method outperforms existing approaches, demonstrates generalization across human and animal categories, and enables high-quality dynamic 4D content generation from ordinary monocular videos.

0 citationsRead paper

Modality Forcing for Scalable Spatial Generation

Jun 11, 2026

This work proposes Modality Forcing, a novel approach for high-quality conditional or joint image–depth generation under the practical constraint of sparse ground-truth depth supervision. By assigning modality-specific noise levels to image and depth inputs and employing dedicated decoders within a single DiT diffusion model, the method enables scalable joint generation without requiring dense depth annotations. To the best of our knowledge, this is the first framework to achieve such capability using only sparse depth supervision, thereby highlighting the potential of image generation as a spatially aware pretraining objective. Leveraging large-scale text-to-image models trained from scratch with 370M–3.3B parameters, the strongest variant sets a new state of the art in monocular depth estimation, reducing the AbsRel error by 57% compared to existing joint generation approaches.

0 citationsRead paper
Recent publications

Latest Papers

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

Jun 11, 2026

Existing image-to-3D methods struggle to simultaneously achieve pixel-aligned fidelity and complete scene geometry. To address this, this work proposes World Tracing—a generative, pixel-aligned geometric representation that predicts, for each input pixel, an ordered stack of multi-layer 3D points in camera space, where the first layer corresponds to the visible surface and subsequent layers reconstruct occluded structures. This approach uniquely unifies complete geometric generation with precise pixel alignment while preserving 2D–3D correspondences, enabling text-driven editing, novel view synthesis, and training-free integration with textured mesh generators. Built upon a World Tracing Diffusion Transformer (WT-DiT), the method employs factorized yet globally attentive mechanisms to denoise multi-layer geometry, trained via pixel-space flow matching and a hybrid noise schedule. Experiments across object, scene, and dynamic datasets demonstrate significant improvements over current depth estimation and image-to-3D generation models in both visible surface reconstruction and holistic geometry completion.

0 citationsRead paper

Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

Jun 11, 2026

This work addresses the challenge of efficiently generating dynamic 4D human representations from monocular or sparse multi-view videos without relying on explicit geometric priors such as skeletons or depth maps. The authors propose a diffusion-based multi-view video generation method conditioned solely on relative camera poses, integrating spatiotemporal information and SE(3) camera geometry through a five-axis positional encoding scheme. A three-stage curriculum learning strategy enables flexible novel-view synthesis and temporal extrapolation. Built upon the Wan 2.1 1.3B text-to-video diffusion model, the approach incorporates extended spatiotemporal RoPE, historical target-view tokens, and multi-view textual conditioning. Evaluated on the DNA-Rendering and ActorsHQ datasets, the method outperforms existing approaches, demonstrates generalization across human and animal categories, and enables high-quality dynamic 4D content generation from ordinary monocular videos.

0 citationsRead paper

Modality Forcing for Scalable Spatial Generation

Jun 11, 2026

This work proposes Modality Forcing, a novel approach for high-quality conditional or joint image–depth generation under the practical constraint of sparse ground-truth depth supervision. By assigning modality-specific noise levels to image and depth inputs and employing dedicated decoders within a single DiT diffusion model, the method enables scalable joint generation without requiring dense depth annotations. To the best of our knowledge, this is the first framework to achieve such capability using only sparse depth supervision, thereby highlighting the potential of image generation as a spatially aware pretraining objective. Leveraging large-scale text-to-image models trained from scratch with 370M–3.3B parameters, the strongest variant sets a new state of the art in monocular depth estimation, reducing the AbsRel error by 57% compared to existing joint generation approaches.

0 citationsRead paper