Institution profile

RMIT International University Vietnam

Academic institutionasia · vn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

VietNormalizer: An Open-Source, Dependency-Free Python Library for Vietnamese Text Normalization in TTS and NLP Applications

Mar 04, 2026

This work addresses the challenge posed by prevalent non-standard textual elements in Vietnamese—such as numerals, dates, currencies, abbreviations, and loanwords—which hinder the effective processing of text-to-speech (TTS) and natural language processing (NLP) systems. Existing normalization tools are often either computationally heavy, incompletely covered, or dependent on external services, making standalone deployment impractical. To overcome these limitations, we propose a lightweight, zero-dependency, rule-based unified text normalization pipeline that integrates precompiled regular expressions, a rule engine, CSV-based dictionary mappings, transliteration algorithms, and Unicode normalization. Notably, the system operates without neural networks or external APIs yet comprehensively handles diverse non-standard tokens. The solution is open-sourced via PyPI and GitHub, supports pip installation, and offers high throughput, low memory footprint, and strong extensibility, providing the first self-contained text normalization framework tailored for low-resource tonal languages like Vietnamese.

0 citationsRead paper

VietSuperSpeech: A Large-Scale Vietnamese Conversational Speech Dataset for ASR Fine-Tuning in Chatbot, Customer Support, and Call Center Applications

Mar 02, 2026

This work addresses the scarcity of natural conversational speech data for Vietnamese automatic speech recognition (ASR), as existing datasets predominantly consist of formal read or news-style utterances. To bridge this gap, we introduce VietSuperSpeech—a dataset comprising 267.39 hours of spontaneous, real-world conversational audio drawn from informal contexts such as casual chats, vlogs, and overseas Vietnamese community interactions. Using the Zipformer-30M-RNNT-6000h model via Sherpa-ONNX, we generate pseudo-labels for the recordings, which are standardized to 16 kHz mono WAV format. After rigorous quality filtering, the data are split into training and test sets using a fixed random seed. The publicly released dataset includes 52,023 annotated utterances, totaling 13.8 million Vietnamese characters with full diacritic annotation, significantly enhancing ASR performance in conversational Vietnamese scenarios.

0 citationsRead paper

Harnessing the Power of Large Language Models for Software Testing Education: A Focus on ISTQB Syllabus

Oct 25, 2025

This study addresses the integration of large language models (LLMs) into higher education for software testing under the ISTQB certification framework. Methodologically, we construct the first standardized ISTQB examination dataset spanning 11 years and comprising 1,145 questions; design domain-specific prompt engineering strategies; and establish a dedicated evaluation framework jointly assessing knowledge alignment and explanation quality. Contributions include the first systematic validation of mainstream LLMs on ISTQB exam preparation tasks—demonstrating high accuracy and explanatory fidelity across multiple models; proposing a reproducible AI-augmented pedagogical integration framework; and open-sourcing both the dataset and evaluation benchmark. These resources constitute foundational infrastructure and an actionable paradigm for AI-driven software engineering education.

0 citationsRead paper
Recent publications

Latest Papers

VietNormalizer: An Open-Source, Dependency-Free Python Library for Vietnamese Text Normalization in TTS and NLP Applications

Mar 04, 2026

This work addresses the challenge posed by prevalent non-standard textual elements in Vietnamese—such as numerals, dates, currencies, abbreviations, and loanwords—which hinder the effective processing of text-to-speech (TTS) and natural language processing (NLP) systems. Existing normalization tools are often either computationally heavy, incompletely covered, or dependent on external services, making standalone deployment impractical. To overcome these limitations, we propose a lightweight, zero-dependency, rule-based unified text normalization pipeline that integrates precompiled regular expressions, a rule engine, CSV-based dictionary mappings, transliteration algorithms, and Unicode normalization. Notably, the system operates without neural networks or external APIs yet comprehensively handles diverse non-standard tokens. The solution is open-sourced via PyPI and GitHub, supports pip installation, and offers high throughput, low memory footprint, and strong extensibility, providing the first self-contained text normalization framework tailored for low-resource tonal languages like Vietnamese.

0 citationsRead paper

VietSuperSpeech: A Large-Scale Vietnamese Conversational Speech Dataset for ASR Fine-Tuning in Chatbot, Customer Support, and Call Center Applications

Mar 02, 2026

This work addresses the scarcity of natural conversational speech data for Vietnamese automatic speech recognition (ASR), as existing datasets predominantly consist of formal read or news-style utterances. To bridge this gap, we introduce VietSuperSpeech—a dataset comprising 267.39 hours of spontaneous, real-world conversational audio drawn from informal contexts such as casual chats, vlogs, and overseas Vietnamese community interactions. Using the Zipformer-30M-RNNT-6000h model via Sherpa-ONNX, we generate pseudo-labels for the recordings, which are standardized to 16 kHz mono WAV format. After rigorous quality filtering, the data are split into training and test sets using a fixed random seed. The publicly released dataset includes 52,023 annotated utterances, totaling 13.8 million Vietnamese characters with full diacritic annotation, significantly enhancing ASR performance in conversational Vietnamese scenarios.

0 citationsRead paper

Harnessing the Power of Large Language Models for Software Testing Education: A Focus on ISTQB Syllabus

Oct 25, 2025

This study addresses the integration of large language models (LLMs) into higher education for software testing under the ISTQB certification framework. Methodologically, we construct the first standardized ISTQB examination dataset spanning 11 years and comprising 1,145 questions; design domain-specific prompt engineering strategies; and establish a dedicated evaluation framework jointly assessing knowledge alignment and explanation quality. Contributions include the first systematic validation of mainstream LLMs on ISTQB exam preparation tasks—demonstrating high accuracy and explanatory fidelity across multiple models; proposing a reproducible AI-augmented pedagogical integration framework; and open-sourcing both the dataset and evaluation benchmark. These resources constitute foundational infrastructure and an actionable paradigm for AI-driven software engineering education.

0 citationsRead paper