Institution profile

Ho Chi Minh City University of Economics

Academic institutionasia · vn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

Jul 30, 2026

This work addresses the performance bottleneck in strict zero-shot image captioning caused by the absence of visual feedback during inference. The authors propose a multi-agent framework that enhances vision-language alignment without retraining existing captioners, leveraging multi-stage alignment scoring and unsupervised consensus distillation. A novel multi-checkpoint visual feedback mechanism is introduced during decoding, accompanied by a lightweight learnable reranking module that integrates TriFuse and MemAttend architectures. The approach further incorporates Borda count–based consensus distillation to refine caption selection. Experimental results demonstrate significant improvements: the method achieves a CIDEr score of 117.6 on the COCO Karpathy test set, surpassing the baseline by 9.6 points, and yields gains of 8.1 and 5.7 on Flickr30k and NoCaps, respectively.

0 citationsRead paper

TinyCNNDeep: Lightweight Attention-Based CNN for EEG Classification of Eye States and Sleep Deprivation

Jun 24, 2026

This study addresses the joint four-class classification problem of sleep state (normal vs. sleep-deprived) and eye condition (eyes open vs. closed) using electroencephalography (EEG). The authors propose transforming multi-channel EEG segments into 224×224 grayscale images and employ a lightweight convolutional neural network enhanced with residual connections and Squeeze-and-Excitation channel attention mechanisms. Using only five EEG channels, this approach achieves efficient recognition by integrating image-based EEG representation with channel attention for the first time. After preprocessing involving Z-score normalization, min-max scaling, and center padding, the method attains an average individual accuracy of 83.69% across 35 subjects—outperforming the strongest baseline by 36.03 percentage points and significantly surpassing established models such as EEGNet.

0 citationsRead paper
Recent publications

Latest Papers

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

Jul 30, 2026

This work addresses the performance bottleneck in strict zero-shot image captioning caused by the absence of visual feedback during inference. The authors propose a multi-agent framework that enhances vision-language alignment without retraining existing captioners, leveraging multi-stage alignment scoring and unsupervised consensus distillation. A novel multi-checkpoint visual feedback mechanism is introduced during decoding, accompanied by a lightweight learnable reranking module that integrates TriFuse and MemAttend architectures. The approach further incorporates Borda count–based consensus distillation to refine caption selection. Experimental results demonstrate significant improvements: the method achieves a CIDEr score of 117.6 on the COCO Karpathy test set, surpassing the baseline by 9.6 points, and yields gains of 8.1 and 5.7 on Flickr30k and NoCaps, respectively.

0 citationsRead paper

TinyCNNDeep: Lightweight Attention-Based CNN for EEG Classification of Eye States and Sleep Deprivation

Jun 24, 2026

This study addresses the joint four-class classification problem of sleep state (normal vs. sleep-deprived) and eye condition (eyes open vs. closed) using electroencephalography (EEG). The authors propose transforming multi-channel EEG segments into 224×224 grayscale images and employ a lightweight convolutional neural network enhanced with residual connections and Squeeze-and-Excitation channel attention mechanisms. Using only five EEG channels, this approach achieves efficient recognition by integrating image-based EEG representation with channel attention for the first time. After preprocessing involving Z-score normalization, min-max scaling, and center padding, the method attains an average individual accuracy of 83.69% across 35 subjects—outperforming the strongest baseline by 36.03 percentage points and significantly surpassing established models such as EEGNet.

0 citationsRead paper