Institution profile

Ho Chi Minh City University of Science

Academic institutionasia · vn
Official website
Research library73linked papers
Opportunities0open roles
Selected work

Representative Papers

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

Aug 13, 2026

This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.

0 citationsRead paper

CoDiR: Confidence-Guided Diffusion Refinement for Semi-Supervised Histopathology Segmentation

Aug 12, 2026

This work addresses the challenges of scarce annotations and unreliable pseudo-labels in ambiguous gland regions for semi-supervised histopathology image segmentation. To this end, we propose a confidence-guided diffusion refinement mechanism built upon the Mean Teacher framework. In regions where the teacher model exhibits low prediction confidence, a conditional diffusion model is introduced to perform structure-aware refinement, while high-confidence predictions are leveraged to formulate a weighted consistency loss for training the student model. The proposed approach substantially enhances pseudo-label quality, achieving mDice scores of 88.09%/89.83% on the GlaS dataset and 89.19%/90.29% on the CRAG dataset using only 10% and 20% labeled data, respectively. Notably, the diffusion refinement module alone contributes a +6.36% mDice improvement, consistently outperforming current state-of-the-art methods.

0 citationsRead paper

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Aug 10, 2026

This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.

0 citationsRead paper
Recent publications

Latest Papers

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

Aug 13, 2026

This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.

0 citationsRead paper

CoDiR: Confidence-Guided Diffusion Refinement for Semi-Supervised Histopathology Segmentation

Aug 12, 2026

This work addresses the challenges of scarce annotations and unreliable pseudo-labels in ambiguous gland regions for semi-supervised histopathology image segmentation. To this end, we propose a confidence-guided diffusion refinement mechanism built upon the Mean Teacher framework. In regions where the teacher model exhibits low prediction confidence, a conditional diffusion model is introduced to perform structure-aware refinement, while high-confidence predictions are leveraged to formulate a weighted consistency loss for training the student model. The proposed approach substantially enhances pseudo-label quality, achieving mDice scores of 88.09%/89.83% on the GlaS dataset and 89.19%/90.29% on the CRAG dataset using only 10% and 20% labeled data, respectively. Notably, the diffusion refinement module alone contributes a +6.36% mDice improvement, consistently outperforming current state-of-the-art methods.

0 citationsRead paper

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Aug 10, 2026

This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.

0 citationsRead paper