Institution profile

DeepGlint

Industry researchasia · cn
Official website
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets

May 28, 2026

This work proposes an efficient method for evaluating the intrinsic quality (IQ) of large-scale face recognition datasets without requiring full model training. By integrating neighborhood consistency scores with the effective rank of the embedding space, the approach establishes a lightweight, validation-free quality assessment framework capable of rapidly predicting downstream recognition performance using proxy models or dataset subsets. Experimental results demonstrate that the proposed IQ metric accurately forecasts model performance across clean, noisy, and mixed-quality datasets, substantially reducing the cost of data diagnosis and filtering. This provides a practical and scalable tool for preprocessing massive face datasets in real-world applications.

0 citationsRead paper

ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

Oct 21, 2025

CLIP’s text encoder is constrained by a 77-token length limit and lacks multilingual support, hindering fine-grained semantic understanding and cross-lingual cross-modal alignment. To address this, we propose an LLM-based text embedder to replace CLIP’s original text encoder, integrated within a progressive alignment framework. This framework synergistically combines curriculum learning, knowledge distillation, and self-distillation regularization to preserve the pretrained visual encoder’s knowledge. Additionally, we introduce instance-level semantic alignment and structural alignment losses to enhance consistency between image and text representations. Crucially, our method operates without modifying the image encoder. Extensive experiments demonstrate substantial improvements in long-text and multilingual image–text retrieval across Flickr30K, MS-COCO Multilingual, and Long-Caption benchmarks, achieving new state-of-the-art performance. The approach exhibits strong generalization capability and deeper semantic comprehension while maintaining architectural compatibility with existing vision-language models.

0 citationsRead paper

UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction

Oct 02, 2025

This work addresses the challenge of robust 3D scene reconstruction from inconsistent multi-view images—characterized by occlusions, motion blur, and low resolution. We propose a two-stage decoupled framework: first, a video diffusion model learns generic scene priors to achieve cross-view consistent inpainting; second, a neural radiance field (NeRF) is reconstructed from the inpainted views. To our knowledge, this is the first approach to integrate video diffusion models into NeRF reconstruction, enabling robust handling of diverse degradations and controllable stylistic reconstruction. Our method comprises three core components: multi-view-to-video conversion, a consistency-aware inpainting network, and joint NeRF optimization—all trained end-to-end. Extensive experiments on both synthetic and real-world datasets demonstrate significant improvements over state-of-the-art methods, particularly under sparse-view and severely degraded conditions, with enhanced reconstruction fidelity and generalization capability.

0 citationsRead paper

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Sep 10, 2025

In text-driven person retrieval, CLIP-based methods suffer from weak fine-grained feature learning and high sensitivity to textual noise, primarily due to the lack of person-centric training data and insufficient global contrastive learning. To address these issues, this work proposes a novel co-optimization paradigm for data and model. First, leveraging context learning in multimodal large language models (MLLMs), we design a denoising data curation pipeline and release WebPerson—a large-scale, person-centric dataset. Second, we introduce GA-DMS, a gradient-attention-guided dual masking framework that adaptively masks noisy tokens and jointly optimizes masked token prediction with image-text contrastive learning. Extensive experiments demonstrate state-of-the-art performance on benchmarks including CUHK-PEDES and RSTPReid, achieving significant gains in retrieval accuracy and robustness against textual noise.

0 citationsRead paper

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training

Aug 13, 2025

Existing face representation pretraining methods face three key challenges: insufficient fine-grained semantic modeling, neglect of facial anatomical spatial structure, and low efficiency in leveraging scarce labeled data. This paper proposes an unsupervised face representation pretraining framework addressing these issues. Its core contributions are: (1) a semantic-aware structured masking strategy to enhance local discriminability; (2) a candidate codebook-driven block-level encoding scheme with patch-pixel alignment, explicitly enforcing facial geometric constraints; and (3) an end-to-end learnable codebook coupled with spatial consistency regularization. Trained solely on 2 million unlabeled face images, the method achieves state-of-the-art performance across diverse downstream tasks—including face recognition, pose estimation, and occlusion-robust analysis—particularly excelling under extreme pose variations, severe occlusions, and challenging illumination conditions.

0 citationsRead paper
Recent publications

Latest Papers

Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets

May 28, 2026

This work proposes an efficient method for evaluating the intrinsic quality (IQ) of large-scale face recognition datasets without requiring full model training. By integrating neighborhood consistency scores with the effective rank of the embedding space, the approach establishes a lightweight, validation-free quality assessment framework capable of rapidly predicting downstream recognition performance using proxy models or dataset subsets. Experimental results demonstrate that the proposed IQ metric accurately forecasts model performance across clean, noisy, and mixed-quality datasets, substantially reducing the cost of data diagnosis and filtering. This provides a practical and scalable tool for preprocessing massive face datasets in real-world applications.

0 citationsRead paper

ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

Oct 21, 2025

CLIP’s text encoder is constrained by a 77-token length limit and lacks multilingual support, hindering fine-grained semantic understanding and cross-lingual cross-modal alignment. To address this, we propose an LLM-based text embedder to replace CLIP’s original text encoder, integrated within a progressive alignment framework. This framework synergistically combines curriculum learning, knowledge distillation, and self-distillation regularization to preserve the pretrained visual encoder’s knowledge. Additionally, we introduce instance-level semantic alignment and structural alignment losses to enhance consistency between image and text representations. Crucially, our method operates without modifying the image encoder. Extensive experiments demonstrate substantial improvements in long-text and multilingual image–text retrieval across Flickr30K, MS-COCO Multilingual, and Long-Caption benchmarks, achieving new state-of-the-art performance. The approach exhibits strong generalization capability and deeper semantic comprehension while maintaining architectural compatibility with existing vision-language models.

0 citationsRead paper

UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction

Oct 02, 2025

This work addresses the challenge of robust 3D scene reconstruction from inconsistent multi-view images—characterized by occlusions, motion blur, and low resolution. We propose a two-stage decoupled framework: first, a video diffusion model learns generic scene priors to achieve cross-view consistent inpainting; second, a neural radiance field (NeRF) is reconstructed from the inpainted views. To our knowledge, this is the first approach to integrate video diffusion models into NeRF reconstruction, enabling robust handling of diverse degradations and controllable stylistic reconstruction. Our method comprises three core components: multi-view-to-video conversion, a consistency-aware inpainting network, and joint NeRF optimization—all trained end-to-end. Extensive experiments on both synthetic and real-world datasets demonstrate significant improvements over state-of-the-art methods, particularly under sparse-view and severely degraded conditions, with enhanced reconstruction fidelity and generalization capability.

0 citationsRead paper

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Sep 10, 2025

In text-driven person retrieval, CLIP-based methods suffer from weak fine-grained feature learning and high sensitivity to textual noise, primarily due to the lack of person-centric training data and insufficient global contrastive learning. To address these issues, this work proposes a novel co-optimization paradigm for data and model. First, leveraging context learning in multimodal large language models (MLLMs), we design a denoising data curation pipeline and release WebPerson—a large-scale, person-centric dataset. Second, we introduce GA-DMS, a gradient-attention-guided dual masking framework that adaptively masks noisy tokens and jointly optimizes masked token prediction with image-text contrastive learning. Extensive experiments demonstrate state-of-the-art performance on benchmarks including CUHK-PEDES and RSTPReid, achieving significant gains in retrieval accuracy and robustness against textual noise.

0 citationsRead paper

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training

Aug 13, 2025

Existing face representation pretraining methods face three key challenges: insufficient fine-grained semantic modeling, neglect of facial anatomical spatial structure, and low efficiency in leveraging scarce labeled data. This paper proposes an unsupervised face representation pretraining framework addressing these issues. Its core contributions are: (1) a semantic-aware structured masking strategy to enhance local discriminability; (2) a candidate codebook-driven block-level encoding scheme with patch-pixel alignment, explicitly enforcing facial geometric constraints; and (3) an end-to-end learnable codebook coupled with spatial consistency regularization. Trained solely on 2 million unlabeled face images, the method achieves state-of-the-art performance across diverse downstream tasks—including face recognition, pose estimation, and occlusion-robust analysis—particularly excelling under extreme pose variations, severe occlusions, and challenging illumination conditions.

0 citationsRead paper