Relative Time Intervals Representation for Word-level Timestamping with Masked Training
本文通过使用相对时间戳和混合微调策略,增强SpeechLLM对语音内容和时间结构的联合建模能力,提高时间戳预测准确性。
本文通过使用相对时间戳和混合微调策略,增强SpeechLLM对语音内容和时间结构的联合建模能力,提高时间戳预测准确性。
This work addresses the challenge of incomparable evaluations and irreproducible results in speech understanding models, which often arise from discrepancies in post-processing, data handling, and pipeline design during deployment-oriented model selection. To this end, the authors propose SURE, a unified experimental framework that enables fair evaluation across diverse paradigms—from conventional pipelines to speech large language models—under realistic acoustic and linguistic stressors. SURE achieves this through standardized prediction formats, consistent normalization strategies, and a unified scoring mechanism. Furthermore, it introduces an agent-assisted training conversion pipeline that automatically maps published code into versioned, executable training workflows. This study presents the first unified and reproducible approach for both evaluating and training speech understanding systems across modeling paradigms, substantially enhancing comparability and reproducibility in real-world deployment scenarios.
This work addresses the scarcity of high-quality audio-video–text paired data and the limited precision of existing alignment methods in post-training speech large language models. The authors propose a controllable CTC-based simulation framework that, for the first time, enables explicit control over word error rate (WER) and uncertainty, generating difficulty-adjustable textual supervision signals without requiring text-to-speech (TTS) synthesis. This approach facilitates principled curriculum learning strategies and achieves significant improvements over strong baselines—including TASU, text-only fine-tuning, and TTS-augmented methods—across multiple domain transfer tasks, while effectively mitigating performance degradation on the source domain.
This work addresses the vulnerability of multimodal large language models to progressive harmful intent attacks in multi-turn dialogues, a challenge inadequately mitigated by existing single-turn alignment methods. To tackle this, the authors introduce InterSafe-V, the first open-source dataset dedicated to multi-turn multimodal safety, comprising 11,270 dialogues. They further propose the AM³Safety framework, which employs a cold-start refusal phase and a turn-aware dual-objective reward mechanism to guide GRPO fine-tuning for efficient and robust safety alignment. Evaluated on Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B, the approach reduces attack success rates by over 10%, improves harmlessness by at least 8%, and enhances helpfulness by more than 13%, all while preserving general capabilities.
Existing NL2GQL research primarily focuses on single-turn translation, failing to address the prevalent multi-turn, context-dependent interactions between users and graph databases, and suffers from a lack of high-quality, multi-turn annotated datasets. To bridge this gap, we propose an LLM-based automated framework for constructing multi-turn NL2GQL data, integrating dialogue context modeling with graph query syntax constraints. Leveraging this method, we introduce MTGQL—the first domain-specific, multi-turn graph query dataset in finance—comprising over 10,000 dialogue turns. Using MTGQL, we design and evaluate three baseline models across multiple dimensions. Experimental results demonstrate the feasibility of multi-turn semantic understanding and GQL generation, thereby filling critical gaps in both data resources and benchmarking infrastructure. Our work establishes a reproducible foundation and methodological paradigm for dynamic graph query understanding.
本文通过使用相对时间戳和混合微调策略,增强SpeechLLM对语音内容和时间结构的联合建模能力,提高时间戳预测准确性。
This work addresses the challenge of incomparable evaluations and irreproducible results in speech understanding models, which often arise from discrepancies in post-processing, data handling, and pipeline design during deployment-oriented model selection. To this end, the authors propose SURE, a unified experimental framework that enables fair evaluation across diverse paradigms—from conventional pipelines to speech large language models—under realistic acoustic and linguistic stressors. SURE achieves this through standardized prediction formats, consistent normalization strategies, and a unified scoring mechanism. Furthermore, it introduces an agent-assisted training conversion pipeline that automatically maps published code into versioned, executable training workflows. This study presents the first unified and reproducible approach for both evaluating and training speech understanding systems across modeling paradigms, substantially enhancing comparability and reproducibility in real-world deployment scenarios.
This work addresses the scarcity of high-quality audio-video–text paired data and the limited precision of existing alignment methods in post-training speech large language models. The authors propose a controllable CTC-based simulation framework that, for the first time, enables explicit control over word error rate (WER) and uncertainty, generating difficulty-adjustable textual supervision signals without requiring text-to-speech (TTS) synthesis. This approach facilitates principled curriculum learning strategies and achieves significant improvements over strong baselines—including TASU, text-only fine-tuning, and TTS-augmented methods—across multiple domain transfer tasks, while effectively mitigating performance degradation on the source domain.
This work addresses the vulnerability of multimodal large language models to progressive harmful intent attacks in multi-turn dialogues, a challenge inadequately mitigated by existing single-turn alignment methods. To tackle this, the authors introduce InterSafe-V, the first open-source dataset dedicated to multi-turn multimodal safety, comprising 11,270 dialogues. They further propose the AM³Safety framework, which employs a cold-start refusal phase and a turn-aware dual-objective reward mechanism to guide GRPO fine-tuning for efficient and robust safety alignment. Evaluated on Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B, the approach reduces attack success rates by over 10%, improves harmlessness by at least 8%, and enhances helpfulness by more than 13%, all while preserving general capabilities.
Existing NL2GQL research primarily focuses on single-turn translation, failing to address the prevalent multi-turn, context-dependent interactions between users and graph databases, and suffers from a lack of high-quality, multi-turn annotated datasets. To bridge this gap, we propose an LLM-based automated framework for constructing multi-turn NL2GQL data, integrating dialogue context modeling with graph query syntax constraints. Leveraging this method, we introduce MTGQL—the first domain-specific, multi-turn graph query dataset in finance—comprising over 10,000 dialogue turns. Using MTGQL, we design and evaluate three baseline models across multiple dimensions. Experimental results demonstrate the feasibility of multi-turn semantic understanding and GQL generation, thereby filling critical gaps in both data resources and benchmarking infrastructure. Our work establishes a reproducible foundation and methodological paradigm for dynamic graph query understanding.