Institution profile

North China University of Technology

Academic institutionasia · cn
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

Aug 10, 2026

Existing methods struggle to accurately generate and edit fine-grained geometric details of 3D faces—such as eyebrow tension or cheek contraction—from long textual descriptions. To address this challenge, this work introduces FaME-G2E, a large-scale multimodal dataset, and proposes RAGMesh, a retrieval-augmented framework that integrates text-guided global and regional geometric priors in blendshape space. The framework innovatively combines a multi-scale retrieval fusion (MSRF) module with an adaptive RAG-guided supervision (AdaRAGS) mechanism to achieve precise semantic alignment and localized deformation control. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in terms of local geometric accuracy, text controllability, regional editing precision, and inference efficiency.

0 citationsRead paper

An AI4AI Framework for Visual Token Pruning

Aug 07, 2026

This work addresses the limitations of existing vision token pruning methods, which rely on handcrafted heuristics and struggle to generalize across diverse model architectures and pruning objectives. To overcome this, the authors propose AutoPrune, a training-free AI4AI framework that leverages large language models (LLMs) to automatically design effective pruning strategies. Key innovations include TPDSL, a domain-specific language tailored for token pruning; a residual representation of search states to emphasize critical strategy components; and structured guidance for LLMs to generate algorithms within a constrained space. Extensive experiments demonstrate AutoPrune’s superior performance across 14 benchmarks and three multimodal large language models: it retains over 99% of original accuracy even after removing 94.4% of visual tokens, achieving a 9.9× reduction in FLOPs and a 6.4× speedup in prefill latency.

0 citationsRead paper
Recent publications

Latest Papers

RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

Aug 10, 2026

Existing methods struggle to accurately generate and edit fine-grained geometric details of 3D faces—such as eyebrow tension or cheek contraction—from long textual descriptions. To address this challenge, this work introduces FaME-G2E, a large-scale multimodal dataset, and proposes RAGMesh, a retrieval-augmented framework that integrates text-guided global and regional geometric priors in blendshape space. The framework innovatively combines a multi-scale retrieval fusion (MSRF) module with an adaptive RAG-guided supervision (AdaRAGS) mechanism to achieve precise semantic alignment and localized deformation control. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in terms of local geometric accuracy, text controllability, regional editing precision, and inference efficiency.

0 citationsRead paper

An AI4AI Framework for Visual Token Pruning

Aug 07, 2026

This work addresses the limitations of existing vision token pruning methods, which rely on handcrafted heuristics and struggle to generalize across diverse model architectures and pruning objectives. To overcome this, the authors propose AutoPrune, a training-free AI4AI framework that leverages large language models (LLMs) to automatically design effective pruning strategies. Key innovations include TPDSL, a domain-specific language tailored for token pruning; a residual representation of search states to emphasize critical strategy components; and structured guidance for LLMs to generate algorithms within a constrained space. Extensive experiments demonstrate AutoPrune’s superior performance across 14 benchmarks and three multimodal large language models: it retains over 99% of original accuracy even after removing 94.4% of visual tokens, achieving a 9.9× reduction in FLOPs and a 6.4× speedup in prefill latency.

0 citationsRead paper