World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
Existing image-to-3D methods struggle to simultaneously achieve pixel-aligned fidelity and complete scene geometry. To address this, this work proposes World Tracing—a generative, pixel-aligned geometric representation that predicts, for each input pixel, an ordered stack of multi-layer 3D points in camera space, where the first layer corresponds to the visible surface and subsequent layers reconstruct occluded structures. This approach uniquely unifies complete geometric generation with precise pixel alignment while preserving 2D–3D correspondences, enabling text-driven editing, novel view synthesis, and training-free integration with textured mesh generators. Built upon a World Tracing Diffusion Transformer (WT-DiT), the method employs factorized yet globally attentive mechanisms to denoise multi-layer geometry, trained via pixel-space flow matching and a hybrid noise schedule. Experiments across object, scene, and dynamic datasets demonstrate significant improvements over current depth estimation and image-to-3D generation models in both visible surface reconstruction and holistic geometry completion.