Institution profile

iQIYI

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

Sep 18, 2025

Low-resource languages like Thai face critical challenges in speech large language modeling (SLLM), including poor speech encoder performance, weak multimodal understanding capabilities, high computational cost of ASR-based forced alignment, and scarcity of paired speech-text data. To address these, this work proposes a systematic solution: (1) the first Thai self-supervised speech encoder, XLSR-Thai; (2) U-Align, a lightweight cross-modal alignment method that replaces expensive ASR-based forced alignment; and (3) Thai-SUP, a scalable Thai understanding data synthesis framework generating over 1,000 hours of high-quality, multitask training data. Through joint optimization via self-supervised pretraining, U-Align fine-tuning, and cross-lingual transfer, our approach significantly improves Thai speech recognition, semantic understanding, and instruction-following performance. All models and datasets are publicly released, establishing essential infrastructure for low-resource speech understanding research.

0 citationsRead paper

Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR

May 28, 2025

To address the dual challenges of scarce high-quality labeled data and high computational overhead in low-resource Thai automatic speech recognition (ASR), this paper proposes EThai-ASR—the first large language model (LLM)-driven efficient Thai ASR system. Methodologically, we design a self-evolving weak-label refinement strategy to enhance speech encoder robustness; introduce a plug-and-play trimodal sequence compression module that enables dynamic-length compression, significantly reducing computation while preserving modeling capacity; and construct an end-to-end architecture comprising a speech encoder, a connector module, and a Thai-specific LLM-based decoder. Evaluated on multiple Thai ASR benchmarks, EThai-ASR achieves state-of-the-art performance. Furthermore, we publicly release a high-quality refined transcription dataset, establishing a new paradigm and foundational resource for low-resource speech recognition.

0 citationsRead paper

MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception

Jan 02, 2025

Evaluating the robustness and safety of autonomous driving multi-sensor fusion models under sensor failures and adverse environmental conditions remains a critical unsolved challenge. Method: This paper introduces MSC-Bench, the first perception-oriented multi-sensor degradation benchmark. Built upon nuScenes and Argoverse, it provides a reproducible, sensor-level degradation injection framework and formally defines 16 fine-grained degradation combinations—spanning single- and dual-modality (camera/LiDAR) degradations—including weather-induced interference and hardware failures. Contribution/Results: We systematically evaluate six 3D object detection models and four high-definition map estimation models, revealing substantial performance degradation under realistic degradations. The complete toolchain—including code, degradation configurations, and pre-trained model checkpoints—is publicly released. MSC-Bench fills a key gap in multimodal robustness evaluation and establishes a standardized, safety-driven assessment foundation for future research on resilient multi-sensor fusion.

0 citationsRead paper
Recent publications

Latest Papers

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

Sep 18, 2025

Low-resource languages like Thai face critical challenges in speech large language modeling (SLLM), including poor speech encoder performance, weak multimodal understanding capabilities, high computational cost of ASR-based forced alignment, and scarcity of paired speech-text data. To address these, this work proposes a systematic solution: (1) the first Thai self-supervised speech encoder, XLSR-Thai; (2) U-Align, a lightweight cross-modal alignment method that replaces expensive ASR-based forced alignment; and (3) Thai-SUP, a scalable Thai understanding data synthesis framework generating over 1,000 hours of high-quality, multitask training data. Through joint optimization via self-supervised pretraining, U-Align fine-tuning, and cross-lingual transfer, our approach significantly improves Thai speech recognition, semantic understanding, and instruction-following performance. All models and datasets are publicly released, establishing essential infrastructure for low-resource speech understanding research.

0 citationsRead paper

Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR

May 28, 2025

To address the dual challenges of scarce high-quality labeled data and high computational overhead in low-resource Thai automatic speech recognition (ASR), this paper proposes EThai-ASR—the first large language model (LLM)-driven efficient Thai ASR system. Methodologically, we design a self-evolving weak-label refinement strategy to enhance speech encoder robustness; introduce a plug-and-play trimodal sequence compression module that enables dynamic-length compression, significantly reducing computation while preserving modeling capacity; and construct an end-to-end architecture comprising a speech encoder, a connector module, and a Thai-specific LLM-based decoder. Evaluated on multiple Thai ASR benchmarks, EThai-ASR achieves state-of-the-art performance. Furthermore, we publicly release a high-quality refined transcription dataset, establishing a new paradigm and foundational resource for low-resource speech recognition.

0 citationsRead paper

MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception

Jan 02, 2025

Evaluating the robustness and safety of autonomous driving multi-sensor fusion models under sensor failures and adverse environmental conditions remains a critical unsolved challenge. Method: This paper introduces MSC-Bench, the first perception-oriented multi-sensor degradation benchmark. Built upon nuScenes and Argoverse, it provides a reproducible, sensor-level degradation injection framework and formally defines 16 fine-grained degradation combinations—spanning single- and dual-modality (camera/LiDAR) degradations—including weather-induced interference and hardware failures. Contribution/Results: We systematically evaluate six 3D object detection models and four high-definition map estimation models, revealing substantial performance degradation under realistic degradations. The complete toolchain—including code, degradation configurations, and pre-trained model checkpoints—is publicly released. MSC-Bench fills a key gap in multimodal robustness evaluation and establishes a standardized, safety-driven assessment foundation for future research on resilient multi-sensor fusion.

0 citationsRead paper