RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing
Existing methods struggle to accurately generate and edit fine-grained geometric details of 3D faces—such as eyebrow tension or cheek contraction—from long textual descriptions. To address this challenge, this work introduces FaME-G2E, a large-scale multimodal dataset, and proposes RAGMesh, a retrieval-augmented framework that integrates text-guided global and regional geometric priors in blendshape space. The framework innovatively combines a multi-scale retrieval fusion (MSRF) module with an adaptive RAG-guided supervision (AdaRAGS) mechanism to achieve precise semantic alignment and localized deformation control. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in terms of local geometric accuracy, text controllability, regional editing precision, and inference efficiency.