Institution profile

AISpeech Ltd

Industry researchasia · cn
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

A Unified and Reproducible Experimentation Framework for Speech Understanding

May 29, 2026

This work addresses the challenge of incomparable evaluations and irreproducible results in speech understanding models, which often arise from discrepancies in post-processing, data handling, and pipeline design during deployment-oriented model selection. To this end, the authors propose SURE, a unified experimental framework that enables fair evaluation across diverse paradigms—from conventional pipelines to speech large language models—under realistic acoustic and linguistic stressors. SURE achieves this through standardized prediction formats, consistent normalization strategies, and a unified scoring mechanism. Furthermore, it introduces an agent-assisted training conversion pipeline that automatically maps published code into versioned, executable training workflows. This study presents the first unified and reproducible approach for both evaluating and training speech understanding systems across modeling paradigms, substantially enhancing comparability and reproducibility in real-world deployment scenarios.

0 citationsRead paper

TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs

Apr 09, 2026

This work addresses the scarcity of high-quality audio-video–text paired data and the limited precision of existing alignment methods in post-training speech large language models. The authors propose a controllable CTC-based simulation framework that, for the first time, enables explicit control over word error rate (WER) and uncertainty, generating difficulty-adjustable textual supervision signals without requiring text-to-speech (TTS) synthesis. This approach facilitates principled curriculum learning strategies and achieves significant improvements over strong baselines—including TASU, text-only fine-tuning, and TTS-augmented methods—across multiple domain transfer tasks, while effectively mitigating performance degradation on the source domain.

0 citationsRead paper

AM3Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs

Jan 08, 2026arXiv.org

This work addresses the vulnerability of multimodal large language models to progressive harmful intent attacks in multi-turn dialogues, a challenge inadequately mitigated by existing single-turn alignment methods. To tackle this, the authors introduce InterSafe-V, the first open-source dataset dedicated to multi-turn multimodal safety, comprising 11,270 dialogues. They further propose the AM³Safety framework, which employs a cold-start refusal phase and a turn-aware dual-objective reward mechanism to guide GRPO fine-tuning for efficient and robust safety alignment. Evaluated on Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B, the approach reduces attack success rates by over 10%, improves harmlessness by at least 8%, and enhances helpfulness by more than 13%, all while preserving general capabilities.

0 citationsRead paper

Multi-turn Natural Language to Graph Query Language Translation

Aug 03, 2025

Existing NL2GQL research primarily focuses on single-turn translation, failing to address the prevalent multi-turn, context-dependent interactions between users and graph databases, and suffers from a lack of high-quality, multi-turn annotated datasets. To bridge this gap, we propose an LLM-based automated framework for constructing multi-turn NL2GQL data, integrating dialogue context modeling with graph query syntax constraints. Leveraging this method, we introduce MTGQL—the first domain-specific, multi-turn graph query dataset in finance—comprising over 10,000 dialogue turns. Using MTGQL, we design and evaluate three baseline models across multiple dimensions. Experimental results demonstrate the feasibility of multi-turn semantic understanding and GQL generation, thereby filling critical gaps in both data resources and benchmarking infrastructure. Our work establishes a reproducible foundation and methodological paradigm for dynamic graph query understanding.

0 citationsRead paper
Recent publications

Latest Papers

A Unified and Reproducible Experimentation Framework for Speech Understanding

May 29, 2026

This work addresses the challenge of incomparable evaluations and irreproducible results in speech understanding models, which often arise from discrepancies in post-processing, data handling, and pipeline design during deployment-oriented model selection. To this end, the authors propose SURE, a unified experimental framework that enables fair evaluation across diverse paradigms—from conventional pipelines to speech large language models—under realistic acoustic and linguistic stressors. SURE achieves this through standardized prediction formats, consistent normalization strategies, and a unified scoring mechanism. Furthermore, it introduces an agent-assisted training conversion pipeline that automatically maps published code into versioned, executable training workflows. This study presents the first unified and reproducible approach for both evaluating and training speech understanding systems across modeling paradigms, substantially enhancing comparability and reproducibility in real-world deployment scenarios.

0 citationsRead paper

TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs

Apr 09, 2026

This work addresses the scarcity of high-quality audio-video–text paired data and the limited precision of existing alignment methods in post-training speech large language models. The authors propose a controllable CTC-based simulation framework that, for the first time, enables explicit control over word error rate (WER) and uncertainty, generating difficulty-adjustable textual supervision signals without requiring text-to-speech (TTS) synthesis. This approach facilitates principled curriculum learning strategies and achieves significant improvements over strong baselines—including TASU, text-only fine-tuning, and TTS-augmented methods—across multiple domain transfer tasks, while effectively mitigating performance degradation on the source domain.

0 citationsRead paper

AM3Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs

Jan 08, 2026arXiv.org

This work addresses the vulnerability of multimodal large language models to progressive harmful intent attacks in multi-turn dialogues, a challenge inadequately mitigated by existing single-turn alignment methods. To tackle this, the authors introduce InterSafe-V, the first open-source dataset dedicated to multi-turn multimodal safety, comprising 11,270 dialogues. They further propose the AM³Safety framework, which employs a cold-start refusal phase and a turn-aware dual-objective reward mechanism to guide GRPO fine-tuning for efficient and robust safety alignment. Evaluated on Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B, the approach reduces attack success rates by over 10%, improves harmlessness by at least 8%, and enhances helpfulness by more than 13%, all while preserving general capabilities.

0 citationsRead paper

Multi-turn Natural Language to Graph Query Language Translation

Aug 03, 2025

Existing NL2GQL research primarily focuses on single-turn translation, failing to address the prevalent multi-turn, context-dependent interactions between users and graph databases, and suffers from a lack of high-quality, multi-turn annotated datasets. To bridge this gap, we propose an LLM-based automated framework for constructing multi-turn NL2GQL data, integrating dialogue context modeling with graph query syntax constraints. Leveraging this method, we introduce MTGQL—the first domain-specific, multi-turn graph query dataset in finance—comprising over 10,000 dialogue turns. Using MTGQL, we design and evaluate three baseline models across multiple dimensions. Experimental results demonstrate the feasibility of multi-turn semantic understanding and GQL generation, thereby filling critical gaps in both data resources and benchmarking infrastructure. Our work establishes a reproducible foundation and methodological paradigm for dynamic graph query understanding.

0 citationsRead paper