TGRHuman: Text-Guided Realistic 3D Human Generation via Diffusion Renderer

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing text-to-3D human generation methods struggle to simultaneously achieve geometric and textural consistency, high-fidelity detail, and computational efficiency. This work proposes a decoupled generation framework that separates geometry and texture synthesis. It constructs detailed human shapes through high-resolution multi-view normal estimation and a geometry carving strategy, while introducing a texture prior-guided diffusion-based rendering pipeline to synthesize spatially consistent and photorealistic surface details. The approach effectively supports loose clothing modeling and outperforms current state-of-the-art methods in both geometric accuracy and textural realism, substantially improving overall generation quality and inference efficiency.
📝 Abstract
Realistic 3D human generation plays a crucial role in many graphics applications. However, current methods still struggle to generate high-quality human geometry and texture while maintaining 3D consistency and inference efficiency. In this work, we address these limitations by introducing TGRHuman, a novel approach for generating realistic 3D humans from text. Our method decouples geometry and texture generation to alleviate the issues commonly encountered in NeRF-based methods. Instead of relying on slow, implicit score-distillation-based optimization, we directly use explicit multi-view observation generation and optimization for efficient 3D synthesis. For geometry generation, we propose a high-resolution generative module for multi-view normals together with a geometry-carving strategy that preserves view consistency and supports loose clothing. For texture generation, we produce spatially consistent RGB observations from densely sampled surrounding views using a carefully designed texture-prior acquisition strategy and a diffusion renderer, enabling detailed human texture synthesis. Experiments show that our method can generate high-quality and consistent 3D human geometry and texture efficiently. TGRHuman outperforms existing text-to-3D human methods in both geometry and texture quality.
Problem

Research questions and friction points this paper is trying to address.

3D human generation
text-to-3D
geometry and texture
3D consistency
inference efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-to-3D
Diffusion Renderer
Geometry-Texture Decoupling
Multi-view Consistency
3D Human Generation
🔎 Similar Papers
No similar papers found.