Institution profile

Uniphore Software Systems

Industry researchasia · in
Official website
Research library27linked papers
Opportunities0open roles
Selected work

Representative Papers

Text-only adaptation in LLM-based ASR through text denoising

Jan 28, 2026

This work addresses the performance degradation commonly observed in large language model (LLM)-based speech recognition systems when adapting solely with in-domain text, a process that often disrupts the alignment between speech and text modalities. To mitigate this issue, the authors propose a lightweight text-denoising adaptation approach that reformulates the audio projection task as a text denoising problem. By training the LLM to reconstruct clean transcripts from noisy textual inputs, the method achieves effective domain adaptation without modifying the model architecture or introducing additional parameters. This strategy preserves cross-modal alignment while significantly improving recognition accuracy, yielding up to a 22.1% relative reduction in word error rate on two benchmark datasets—substantially outperforming current state-of-the-art text-only adaptation techniques.

1 citationsRead paper

Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection

Jan 28, 2026

This work addresses the instability and high sensitivity to prompt selection in existing large language model (LLM)-based speech recognition approaches that rely on fixed, manually crafted prompts. To overcome this limitation, the authors propose a model-agnostic, learnable prompt projection module that adaptively maps prompt embeddings into more effective regions of the LLM’s input space, without modifying the underlying LLM architecture. This approach significantly reduces prompt sensitivity and enhances recognition robustness and consistency. Experimental results across four benchmark datasets demonstrate that the proposed method not only consistently outperforms the best handcrafted prompts but also substantially mitigates performance variance, yielding more reliable and stable recognition outcomes.

1 citationsRead paper

LLMs Get Smarter from Targeted Synthetic Multilingual Data

Aug 16, 2026

This study addresses the challenges of insufficient cross-lingual semantic alignment and capability disparities in large language models by proposing Hotfixr, a data-centric framework. This approach precisely identifies multilingual vulnerabilities and generates targeted synthetic data for fine-tuning, thereby effectively enhancing cross-lingual reasoning capabilities and robustness. Experimental results demonstrate that Hotfixr improves in-distribution performance by 6.2% and out-of-distribution language task accuracy by 7.1%, while reducing catastrophic forgetting by 3.7%. These findings indicate significant improvements in generalization and stability for large language models operating in low-resource scenarios, offering a viable solution to mitigate performance gaps across diverse linguistic contexts through strategic data augmentation and optimization.

0 citationsRead paper

When Synthetic Speech Is All You Have: Better Call GRPO

Jul 09, 2026

This study addresses the challenge of adapting large language model–driven automatic speech recognition (ASR) systems in regulated domains such as banking, where access to real-world speech data is severely restricted due to privacy and compliance constraints. To overcome the acoustic distribution mismatch between synthetic and real speech, the authors propose the first application of the reward-free reinforcement learning algorithm GRPO (Group Relative Policy Optimization) for ASR adaptation using only synthetic data. By optimizing policies through rewards based on low word error rate (WER), GRPO reduces WER from 36.71% to 22.09%—a 40% relative improvement over supervised fine-tuning (SFT). Combining SFT with GRPO yields a further 45% reduction. The performance gain is attributed to behavioral policy optimization rather than changes in representation, thereby surpassing the limitations of conventional SFT.

0 citationsRead paper
Recent publications

Latest Papers

LLMs Get Smarter from Targeted Synthetic Multilingual Data

Aug 16, 2026

This study addresses the challenges of insufficient cross-lingual semantic alignment and capability disparities in large language models by proposing Hotfixr, a data-centric framework. This approach precisely identifies multilingual vulnerabilities and generates targeted synthetic data for fine-tuning, thereby effectively enhancing cross-lingual reasoning capabilities and robustness. Experimental results demonstrate that Hotfixr improves in-distribution performance by 6.2% and out-of-distribution language task accuracy by 7.1%, while reducing catastrophic forgetting by 3.7%. These findings indicate significant improvements in generalization and stability for large language models operating in low-resource scenarios, offering a viable solution to mitigate performance gaps across diverse linguistic contexts through strategic data augmentation and optimization.

0 citationsRead paper

When Synthetic Speech Is All You Have: Better Call GRPO

Jul 09, 2026

This study addresses the challenge of adapting large language model–driven automatic speech recognition (ASR) systems in regulated domains such as banking, where access to real-world speech data is severely restricted due to privacy and compliance constraints. To overcome the acoustic distribution mismatch between synthetic and real speech, the authors propose the first application of the reward-free reinforcement learning algorithm GRPO (Group Relative Policy Optimization) for ASR adaptation using only synthetic data. By optimizing policies through rewards based on low word error rate (WER), GRPO reduces WER from 36.71% to 22.09%—a 40% relative improvement over supervised fine-tuning (SFT). Combining SFT with GRPO yields a further 45% reduction. The performance gain is attributed to behavioral policy optimization rather than changes in representation, thereby surpassing the limitations of conventional SFT.

0 citationsRead paper

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

Jun 27, 2026

In privacy-sensitive domains, the scarcity of real speech data and the distributional gap between synthetic and real speech hinder the effective use of synthetic data in automatic speech recognition (ASR). This work addresses this challenge within the SLAM-ASR framework by revealing, for the first time, that discriminative signals distinguishing real from synthetic speech in large language model (LLM) backbones are predominantly localized in early-to-mid layers. Leveraging this insight, the authors propose a synergistic strategy combining a layer selection module with room impulse response (RIR) augmentation. This approach substantially narrows the distributional gap, achieving performance on par with a full real-data baseline using only 25% of real speech (13.6 hours) and even surpassing it at higher proportions, thereby significantly reducing reliance on real speech data.

0 citationsRead paper

When Vision Speaks for Sound

May 13, 2026

This work addresses the "Clever Hans effect" in existing audio-visual multimodal large language models, which often hallucinate audio based on visual cues rather than genuinely comprehending auditory content. To systematically probe whether models achieve authentic audio-visual alignment, the authors propose the Thud framework, introducing three counterfactual audio-editing interventions—Shift, Mute, and Swap—for the first time. They further design a two-stage alignment training strategy that integrates preference pairs derived from these interventions with event-level video preference regularization to enhance audio grounding. Evaluated on a 10K-sample training set, the model demonstrates a 28-percentage-point average accuracy improvement across the three intervention dimensions and achieves consistent, albeit modest, performance gains on standard video and audio-visual question-answering benchmarks.

0 citationsRead paper