Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text
研究通过结合传统词典、未标注文本及预训练字符编码器,解决了西夏文在资源极度稀缺下的分词问题,达到了约0.91的F1分数。
研究通过结合传统词典、未标注文本及预训练字符编码器,解决了西夏文在资源极度稀缺下的分词问题,达到了约0.91的F1分数。
To address the joint modeling challenge of speaker diarization under overlapping speech and low-SNR ASR in the MISP 2025 Challenge, this work proposes an adaptive hybrid diarization architecture and ASR-aware observation enhancement. First, we introduce a novel overlap-adaptive hybrid diarization framework integrating end-to-end segmentation (WavLM), traditional clustering (AHC/i-vector), and guided source separation (GSS). Second, we design an ASR-aware feature compensation mechanism to overcome GSS performance degradation in noisy conditions. Third, we construct an end-to-end and modularly coordinated SD-ASR cascaded system. Our approach achieves first place in both Track 2 (character error rate: 9.48%) and Track 3 (cpCER: 11.56%), demonstrating state-of-the-art robustness and effectiveness in realistic meeting scenarios with overlapping speech and low signal-to-noise ratios.
研究通过结合传统词典、未标注文本及预训练字符编码器,解决了西夏文在资源极度稀缺下的分词问题,达到了约0.91的F1分数。
To address the joint modeling challenge of speaker diarization under overlapping speech and low-SNR ASR in the MISP 2025 Challenge, this work proposes an adaptive hybrid diarization architecture and ASR-aware observation enhancement. First, we introduce a novel overlap-adaptive hybrid diarization framework integrating end-to-end segmentation (WavLM), traditional clustering (AHC/i-vector), and guided source separation (GSS). Second, we design an ASR-aware feature compensation mechanism to overcome GSS performance degradation in noisy conditions. Third, we construct an end-to-end and modularly coordinated SD-ASR cascaded system. Our approach achieves first place in both Track 2 (character error rate: 9.48%) and Track 3 (cpCER: 11.56%), demonstrating state-of-the-art robustness and effectiveness in realistic meeting scenarios with overlapping speech and low signal-to-noise ratios.