Institution profile

Vietnam National University

Academic institutionasia · vn
Official website
Research library222linked papers
Opportunities0open roles
Selected work

Representative Papers

A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks

Jun 04, 2025International Journal of Advanced Computer Science and Applications

This study addresses three critical challenges in medical image segmentation (MIS): weak interpretability of deep neural network (DNN) models, fragmented evaluation frameworks, and insufficient clinical trustworthiness. To tackle these, we systematically introduce the DIKIW (Data–Information–Knowledge–Intelligence–Wisdom) hierarchy into MIS evaluation for the first time. We propose an eXplainable AI (XAI)-driven “Intelligence-to-Wisdom” evolution pathway, integrating multi-level semantic modeling with clinically grounded assessment to establish a comprehensive DIKIW-aligned method taxonomy. Key bottlenecks—including black-box decision-making and inadequate early-lesion representation—are identified, and deployable transparency-enhancing solutions are provided. Our framework significantly improves both interpretability and early detection rates for cancerous lesions, offering theoretical foundations and practical guidelines for high-assurance clinical AI systems.

8 citationsRead paper

ViTextVQA: A Large-Scale Visual Question Answering Dataset for Evaluating Vietnamese Text Comprehension in Images

Apr 16, 2024arXiv.org

This work addresses the lack of evaluation benchmarks for image-text reading comprehension in Vietnamese Visual Question Answering (VQA). We introduce ViTextVQA, the first large-scale Vietnamese VQA dataset explicitly designed for image-text understanding, comprising over 16,000 images and 50,000 OCR-annotated question-answer pairs. We formally define and systematically evaluate image-text comprehension capability in Vietnamese scenes, revealing that OCR token ordering critically affects answer generation. To address this, we propose a dedicated VQA framework integrating OCR sequence modeling (BERT/LSTM), multimodal feature extraction (ViT/CLIP), and cross-modal attention mechanisms. Experiments demonstrate substantial accuracy improvements over mainstream models on Vietnamese image-text understanding tasks. The ViTextVQA dataset is publicly released, establishing a foundational resource for multimodal understanding research in low-resource languages.

3 citationsRead paper

Quantifying Statistical Significance in Diffusion-Based Anomaly Localization via Selective Inference

Feb 19, 2024

Image anomaly localization is critical in medical diagnosis and industrial inspection, yet existing generative-model-based approaches—particularly diffusion models—lack statistical reliability, suffer from model bias and uncertainty, and fail to quantify false-positive risk. This paper introduces selective inference to diffusion-based anomaly localization for the first time, establishing an interpretable statistical inference framework: for each pixel or region in the model-reconstructed image, it performs conditional hypothesis testing and outputs rigorously calibrated p-values to quantify the false-positive probability. Unlike conventional methods lacking theoretical guarantees, our approach enables statistically controlled, significance-aware anomaly localization. Experiments on multiple medical and industrial datasets demonstrate substantial improvements in false-positive rate control, delivering trustworthy, statistically grounded anomaly localization outputs suitable for high-stakes applications.

2 citationsRead paper

Agentic Design Patterns: A System-Theoretic Framework

Jan 27, 2026

Current agent system designs often lack grounding in systems theory, resulting in ad hoc architectures prone to hallucination and reasoning flaws that undermine reliability. This work addresses this gap by introducing systems theory into agent architecture design for the first time, proposing a structured framework composed of five core functional subsystems. Building on this foundation, the authors abstract twelve reusable and clearly categorized agent design patterns. Through the reconstruction and validation of representative frameworks such as ReAct, the proposed approach effectively rectifies inherent architectural deficiencies, significantly enhancing modularity, interpretability, and reliability. This contribution establishes a standardized language and a structured development paradigm for agent engineering, offering a principled foundation for future research and practice.

1 citationsRead paper

Efficient 3D Brain Tumor Segmentation with Axial-Coronal-Sagittal Embedding

May 31, 2025Pacific-Rim Symposium on Image and Video Technology

To address the high computational cost and low pretraining weight utilization of nnU-Net in 3D brain tumor segmentation, this work proposes an efficient, lightweight framework. First, we introduce a novel tri-planar (axial/coronal/sagittal) embedded convolutional architecture to enhance spatial contextual modeling. Second, we design two 2D→3D pretraining weight transfer strategies: one leveraging ImageNet-pretrained generic features, and another exploiting glioma grading task-specific encoder representations. Third, we formulate a joint classification-segmentation learning framework to improve robustness for small tumor subregions. Experiments on the BraTS dataset demonstrate that our method reduces training time by 40% and decreases trainable parameters by 32%, while achieving single-model performance comparable to—or even surpassing—that of conventional cross-validation ensemble models. This advancement significantly improves both efficiency and clinical practicality in medical image segmentation.

1 citationsRead paper
Recent publications

Latest Papers