Institution profile

DGene

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

CityGo: Lightweight Urban Modeling and Rendering with Proxy Buildings and Residual Gaussians

May 27, 2025

To address severe occlusion, geometric incompleteness, high memory overhead, and poor edge-deployment capability in large-scale urban aerial reconstruction, this paper proposes a hybrid representation framework combining proxy building meshes with residual 3D Gaussians. Our method innovatively integrates multi-view stereo (MVS)-derived proxy geometry with depth-guided residual Gaussians, augmented by importance-aware downsampling and joint optimization. We further incorporate zero-order spherical harmonic lighting, image reprojection constraints, and a mobile-GPU-oriented lightweight design. Evaluated on real-world aerial datasets, our approach achieves a 1.4× training speedup while significantly reducing GPU memory consumption and energy usage. Notably, it enables the first real-time rasterization-based rendering of complex urban scenes on consumer-grade mobile GPUs—overcoming fundamental limitations of 3D Gaussian splatting in dense modeling fidelity, prolonged training duration, and on-device adaptability.

0 citationsRead paper

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

Mar 15, 2025

Existing human-centric volumetric video methods are largely confined to dynamic scene replay or character animation, lacking high-fidelity reenactment capability for general dynamic scenes. To address this, we propose the first human-centric volumetric video framework enabling unified “replay → reenactment” modeling. Our approach introduces a hierarchical, disentangled Gaussian representation for motion and appearance, augmented by a semantic-aware alignment module and a deformation-transfer-based motion retargeting mechanism. Integrating Gaussian splatting, Morton encoding, a 2D position-to-attribute mapping CNN, and canonical-space modeling, our method achieves efficient multi-view reconstruction and photorealistic novel-pose rendering. Extensive evaluations on standard benchmarks demonstrate comprehensive superiority over state-of-the-art methods, establishing new paradigmatic benchmarks in reconstruction accuracy, reenactment fidelity, and generalization capability.

0 citationsRead paper
Recent publications

Latest Papers

CityGo: Lightweight Urban Modeling and Rendering with Proxy Buildings and Residual Gaussians

May 27, 2025

To address severe occlusion, geometric incompleteness, high memory overhead, and poor edge-deployment capability in large-scale urban aerial reconstruction, this paper proposes a hybrid representation framework combining proxy building meshes with residual 3D Gaussians. Our method innovatively integrates multi-view stereo (MVS)-derived proxy geometry with depth-guided residual Gaussians, augmented by importance-aware downsampling and joint optimization. We further incorporate zero-order spherical harmonic lighting, image reprojection constraints, and a mobile-GPU-oriented lightweight design. Evaluated on real-world aerial datasets, our approach achieves a 1.4× training speedup while significantly reducing GPU memory consumption and energy usage. Notably, it enables the first real-time rasterization-based rendering of complex urban scenes on consumer-grade mobile GPUs—overcoming fundamental limitations of 3D Gaussian splatting in dense modeling fidelity, prolonged training duration, and on-device adaptability.

0 citationsRead paper

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

Mar 15, 2025

Existing human-centric volumetric video methods are largely confined to dynamic scene replay or character animation, lacking high-fidelity reenactment capability for general dynamic scenes. To address this, we propose the first human-centric volumetric video framework enabling unified “replay → reenactment” modeling. Our approach introduces a hierarchical, disentangled Gaussian representation for motion and appearance, augmented by a semantic-aware alignment module and a deformation-transfer-based motion retargeting mechanism. Integrating Gaussian splatting, Morton encoding, a 2D position-to-attribute mapping CNN, and canonical-space modeling, our method achieves efficient multi-view reconstruction and photorealistic novel-pose rendering. Extensive evaluations on standard benchmarks demonstrate comprehensive superiority over state-of-the-art methods, establishing new paradigmatic benchmarks in reconstruction accuracy, reenactment fidelity, and generalization capability.

0 citationsRead paper