T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition
研究针对台语自动语音识别中变调问题,提出T-SANDHI模型,通过解耦表面声学和词典意图,利用多任务学习和动态门控方法有效解决音调混淆问题。
研究针对台语自动语音识别中变调问题,提出T-SANDHI模型,通过解耦表面声学和词典意图,利用多任务学习和动态门控方法有效解决音调混淆问题。
Existing automatic evaluation methods for text-to-music generation systems struggle to optimize ranking metrics and exhibit weak cross-modal consistency. This work proposes DeRA-MOS, a novel framework that decouples listwise ranking from modality alignment objectives for the first time. It employs a batch-aware listwise ranking loss to optimize the ranking performance of musical impressions and integrates a score-anchored modality alignment loss to enhance semantic consistency between text and music. By explicitly addressing pointwise training bias and modality drift, the proposed approach significantly improves Spearman rank correlation on the MusicEval benchmark, establishing a new paradigm for large-scale evaluation of text-to-music generation systems.
This work addresses the vulnerability of existing automatic audio quality assessment models to data scarcity, which often leads them to learn spurious acoustic correlations tied to specific datasets rather than genuine perceptual quality characteristics. To mitigate this, the authors propose a disentanglement framework based on Domain-Adversarial Training (DAT), systematically exploring diverse domain definition strategies—from explicit metadata to implicit clustering—to separate quality-relevant features from confounding factors. A key insight is that no universally optimal domain partitioning exists; instead, the choice of strategy should be adaptively tailored to different Mean Opinion Score (MOS) dimensions. Experimental results demonstrate that the proposed approach significantly improves correlation with human ratings and exhibits superior generalization to unseen audio generation scenarios.
This work proposes HyperRAG, a novel retrieval-augmented generation (RAG) framework that addresses the limitations of traditional binary knowledge graph–based approaches in multi-hop question answering—namely, rigid retrieval, high computational cost, and insufficient relational expressiveness. HyperRAG is the first to incorporate n-ary hypergraphs into RAG, modeling high-order relational facts through hypergraph structures. It integrates structural-semantic joint reasoning with parametric memory from large language models, featuring a HyperRetriever module that adaptively constructs multi-hop reasoning paths and a HyperMemory mechanism that dynamically guides path expansion. Evaluated on benchmarks including WikiTopics and HotpotQA, HyperRAG significantly outperforms existing methods, achieving an average improvement of 2.95% in MRR and 1.23% in Hits@10, while also demonstrating enhanced interpretability and cross-domain generalization capabilities.
Existing offline signature verification (OSV) methods overly rely on global feature matching, neglecting fine-grained discriminative cues, while standard Transformer backbones tend to weaken local structural patterns, limiting both discriminability and interpretability. To address these issues, we propose DetailSemNet: a novel architecture featuring a detail-semantic fusion module that enables feature disentanglement and re-coupling—preserving stroke-level details while enhancing semantic representation; and a local structural matching mechanism that explicitly models geometric deformation and consistency within critical signature regions. Crucially, our method augments fine-grained modeling capability without modifying the underlying Transformer backbone. Extensive experiments demonstrate state-of-the-art performance on major OSV benchmarks—including CEDAR, Bengali, and Chinese—and significantly superior cross-dataset generalization compared to existing approaches.
研究针对台语自动语音识别中变调问题,提出T-SANDHI模型,通过解耦表面声学和词典意图,利用多任务学习和动态门控方法有效解决音调混淆问题。
Existing automatic evaluation methods for text-to-music generation systems struggle to optimize ranking metrics and exhibit weak cross-modal consistency. This work proposes DeRA-MOS, a novel framework that decouples listwise ranking from modality alignment objectives for the first time. It employs a batch-aware listwise ranking loss to optimize the ranking performance of musical impressions and integrates a score-anchored modality alignment loss to enhance semantic consistency between text and music. By explicitly addressing pointwise training bias and modality drift, the proposed approach significantly improves Spearman rank correlation on the MusicEval benchmark, establishing a new paradigm for large-scale evaluation of text-to-music generation systems.
This work addresses the vulnerability of existing automatic audio quality assessment models to data scarcity, which often leads them to learn spurious acoustic correlations tied to specific datasets rather than genuine perceptual quality characteristics. To mitigate this, the authors propose a disentanglement framework based on Domain-Adversarial Training (DAT), systematically exploring diverse domain definition strategies—from explicit metadata to implicit clustering—to separate quality-relevant features from confounding factors. A key insight is that no universally optimal domain partitioning exists; instead, the choice of strategy should be adaptively tailored to different Mean Opinion Score (MOS) dimensions. Experimental results demonstrate that the proposed approach significantly improves correlation with human ratings and exhibits superior generalization to unseen audio generation scenarios.
This work proposes HyperRAG, a novel retrieval-augmented generation (RAG) framework that addresses the limitations of traditional binary knowledge graph–based approaches in multi-hop question answering—namely, rigid retrieval, high computational cost, and insufficient relational expressiveness. HyperRAG is the first to incorporate n-ary hypergraphs into RAG, modeling high-order relational facts through hypergraph structures. It integrates structural-semantic joint reasoning with parametric memory from large language models, featuring a HyperRetriever module that adaptively constructs multi-hop reasoning paths and a HyperMemory mechanism that dynamically guides path expansion. Evaluated on benchmarks including WikiTopics and HotpotQA, HyperRAG significantly outperforms existing methods, achieving an average improvement of 2.95% in MRR and 1.23% in Hits@10, while also demonstrating enhanced interpretability and cross-domain generalization capabilities.
Existing offline signature verification (OSV) methods overly rely on global feature matching, neglecting fine-grained discriminative cues, while standard Transformer backbones tend to weaken local structural patterns, limiting both discriminability and interpretability. To address these issues, we propose DetailSemNet: a novel architecture featuring a detail-semantic fusion module that enables feature disentanglement and re-coupling—preserving stroke-level details while enhancing semantic representation; and a local structural matching mechanism that explicitly models geometric deformation and consistency within critical signature regions. Crucially, our method augments fine-grained modeling capability without modifying the underlying Transformer backbone. Extensive experiments demonstrate state-of-the-art performance on major OSV benchmarks—including CEDAR, Bengali, and Chinese—and significantly superior cross-dataset generalization compared to existing approaches.