Institution profile

Knovel Engineering

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR

Mar 17, 2026

This work addresses the challenge of developing an efficient and low-cost automatic speech recognition (ASR) system for Singapore’s multilingual setting—encompassing English, Mandarin, Malay, and Tamil—by proposing a balanced sampling fine-tuning strategy that operates without explicit language labels. The approach enables end-to-end training of Qwen3-ASR-0.6B/1.7B models to perform implicit language identification and transcription jointly. The resulting compact multilingual ASR system, Polyglot-Lion-1.7B, achieves an average word error rate of 14.85% across 12 benchmarks, with a remarkably low training cost of just \$81 on a single GPU and an inference speed of 0.10 seconds per sample—approximately 20 times faster than MERaLiON. Despite its significantly reduced model size and computational overhead, the system approaches the performance of large, specialized ASR systems.

0 citationsRead paper

SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging

Apr 14, 2025

Existing medical vision-language models (VLMs) rely on textual instructions, limiting adaptability to real-world clinical settings such as surgery, and lack interpretable reasoning, undermining clinical trustworthiness. This work introduces VoiceMed-VLM—the first end-to-end speech-driven medical VLM—enabling physicians to perform medical image abnormality detection via natural spoken interaction while concurrently generating multi-step, structured clinical reasoning chains. Key contributions include: (1) establishing the first speech-interactive medical VLM paradigm; (2) constructing the first medical reasoning annotation dataset tailored for abnormality detection; and (3) designing an explainability-aligned loss function that jointly optimizes diagnostic outcomes and reasoning process fidelity. VoiceMed-VLM integrates Whisper for speech recognition, fine-tuned Qwen-VL for multimodal understanding, and a dedicated structured reasoning generation module. Evaluated on radiological imaging, it achieves 92.3% abnormality localization accuracy and 86.7% reasoning chain faithfulness, significantly enhancing clinical utility and decision interpretability.

0 citationsRead paper
Recent publications

Latest Papers

Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR

Mar 17, 2026

This work addresses the challenge of developing an efficient and low-cost automatic speech recognition (ASR) system for Singapore’s multilingual setting—encompassing English, Mandarin, Malay, and Tamil—by proposing a balanced sampling fine-tuning strategy that operates without explicit language labels. The approach enables end-to-end training of Qwen3-ASR-0.6B/1.7B models to perform implicit language identification and transcription jointly. The resulting compact multilingual ASR system, Polyglot-Lion-1.7B, achieves an average word error rate of 14.85% across 12 benchmarks, with a remarkably low training cost of just \$81 on a single GPU and an inference speed of 0.10 seconds per sample—approximately 20 times faster than MERaLiON. Despite its significantly reduced model size and computational overhead, the system approaches the performance of large, specialized ASR systems.

0 citationsRead paper

SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging

Apr 14, 2025

Existing medical vision-language models (VLMs) rely on textual instructions, limiting adaptability to real-world clinical settings such as surgery, and lack interpretable reasoning, undermining clinical trustworthiness. This work introduces VoiceMed-VLM—the first end-to-end speech-driven medical VLM—enabling physicians to perform medical image abnormality detection via natural spoken interaction while concurrently generating multi-step, structured clinical reasoning chains. Key contributions include: (1) establishing the first speech-interactive medical VLM paradigm; (2) constructing the first medical reasoning annotation dataset tailored for abnormality detection; and (3) designing an explainability-aligned loss function that jointly optimizes diagnostic outcomes and reasoning process fidelity. VoiceMed-VLM integrates Whisper for speech recognition, fine-tuned Qwen-VL for multimodal understanding, and a dedicated structured reasoning generation module. Evaluated on radiological imaging, it achieves 92.3% abnormality localization accuracy and 86.7% reasoning chain faithfulness, significantly enhancing clinical utility and decision interpretability.

0 citationsRead paper