Institution profile

Korea Culture Technology Institute

Academic institutionasia · kr
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

May 17, 2026

This study investigates the propagation of automatic speech recognition (ASR) errors in Korean spoken question-answering systems that employ an ASR–large language model (LLM) cascade, and the resulting semantic failures. Through ASR error analysis, semantic failure evaluation, and comparative experiments with end-to-end audio-language models, the authors demonstrate that even single-character ASR errors can cause complete downstream QA failure. They find that information loss during ASR is the primary driver of performance degradation, and LLMs of varying capabilities exhibit similar sensitivity to such errors. The results indicate that end-to-end models directly processing audio inputs significantly outperform conventional cascaded architectures in noisy conditions, effectively mitigating semantic information loss caused by transcription errors.

0 citationsRead paper

DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality

Aug 26, 2025

To address the challenges of poor speech recognition for elderly users’ disfluent speech and the inability of conventional ASR-LLM cascaded systems to detect non-speech events (e.g., falls, cries for help), this paper proposes DESAMO—the first edge-deployed, embedded Audio Large Language Model (Audio LLM) system tailored for elderly-friendly smart homes. DESAMO eliminates reliance on ASR by performing multi-granularity audio understanding directly on-device from raw waveforms, jointly modeling natural speech and critical non-speech events while ensuring real-time responsiveness, robustness, and on-device privacy preservation. Experiments demonstrate significant improvements in both elderly speech recognition accuracy and emergency event detection reliability, with average inference latency under 200 ms and end-to-end local data processing. Its core contributions are: (1) the first efficient deployment of an Audio LLM on resource-constrained embedded hardware, and (2) a novel end-to-end audio semantic understanding paradigm specifically designed for elderly users.

0 citationsRead paper
Recent publications

Latest Papers

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

May 17, 2026

This study investigates the propagation of automatic speech recognition (ASR) errors in Korean spoken question-answering systems that employ an ASR–large language model (LLM) cascade, and the resulting semantic failures. Through ASR error analysis, semantic failure evaluation, and comparative experiments with end-to-end audio-language models, the authors demonstrate that even single-character ASR errors can cause complete downstream QA failure. They find that information loss during ASR is the primary driver of performance degradation, and LLMs of varying capabilities exhibit similar sensitivity to such errors. The results indicate that end-to-end models directly processing audio inputs significantly outperform conventional cascaded architectures in noisy conditions, effectively mitigating semantic information loss caused by transcription errors.

0 citationsRead paper

DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality

Aug 26, 2025

To address the challenges of poor speech recognition for elderly users’ disfluent speech and the inability of conventional ASR-LLM cascaded systems to detect non-speech events (e.g., falls, cries for help), this paper proposes DESAMO—the first edge-deployed, embedded Audio Large Language Model (Audio LLM) system tailored for elderly-friendly smart homes. DESAMO eliminates reliance on ASR by performing multi-granularity audio understanding directly on-device from raw waveforms, jointly modeling natural speech and critical non-speech events while ensuring real-time responsiveness, robustness, and on-device privacy preservation. Experiments demonstrate significant improvements in both elderly speech recognition accuracy and emergency event detection reliability, with average inference latency under 200 ms and end-to-end local data processing. Its core contributions are: (1) the first efficient deployment of an Audio LLM on resource-constrained embedded hardware, and (2) a novel end-to-end audio semantic understanding paradigm specifically designed for elderly users.

0 citationsRead paper