The VoiceMOS Challenge 2026: Evaluating Speech Enhancement, Emotional TTS and Accented TTS Systems
该研究通过组织三个不同赛道,评估了语音增强、情感TTS及口音TTS系统的主观语音质量自动预测方法,吸引了18个团队参与并超越基线。
该研究通过组织三个不同赛道,评估了语音增强、情感TTS及口音TTS系统的主观语音质量自动预测方法,吸引了18个团队参与并超越基线。
This work addresses the challenge that conventional time-invariant Koopman operators struggle to effectively model frequency-dependent dynamics in non-stationary time series. To overcome this limitation, the authors propose NDKoop, an end-to-end neural architecture that, for the first time, unifies learnable signal decomposition and Koopman-based dynamic modeling within a single framework. By explicitly separating the trend component (frequency-independent) from periodic components (frequency-dependent) and modeling each with dedicated Koopman networks, NDKoop achieves a more accurate representation of non-stationary dynamics. Extensive experiments demonstrate that NDKoop significantly outperforms existing methods across multiple time series forecasting benchmarks, confirming the efficacy and superiority of the proposed decomposition strategy in scenarios where ideal linearization assumptions do not hold.
This work addresses the limited generalization of existing model inversion attacks, which struggle to recover precise identity information from high-dimensional data. The authors propose GradLock, a training-time injection-based attack that stealthily embeds sensitive training samples into the AI supply chain and employs a dynamic gradient-locking mechanism to prevent payload degradation during optimization, enabling pixel-perfect reconstruction of original data without access to the training environment. GradLock achieves near-instantaneous extraction (<1.0 second), exhibits strong robustness against common deployment optimizations—including quantization, pruning, and fine-tuning—and leverages a stateless deterministic indexing scheme to construct an isolated data vault ensuring payload integrity. Experiments on MNIST, Imagenette, and CelebA demonstrate near-perfect reconstruction (SSIM ≈ 1.0), with 93.3% of users failing to detect the malicious logic, thereby exposing a critical security blind spot in the AI supply chain.
This study addresses foreign language anxiety (FLA), a significant barrier to second language acquisition—particularly in conversational contexts—exacerbated by dialogue agents that often produce overly complex utterances. To mitigate this, the authors propose a multi-agent embodied dialogue system that dynamically adapts linguistic complexity through an innovative “generate–evaluate–regenerate” loop, integrating CEFR proficiency levels with a level classifier to simplify output in real time. This approach uniquely combines multi-agent collaboration with CEFR-aligned simplification, enhancing both comprehensibility and psychological comfort. In a small-scale experiment with Japanese university students, 87.4% of generated utterances fell within ±1 CEFR level of learners’ self-assessed proficiency, substantially outperforming an unsimplified baseline (54.1%), thereby demonstrating the system’s efficacy in tailoring dialogue difficulty to individual learners.
This work addresses the challenge of balancing upload latency and cellular cost in MPQUIC scheduling over heterogeneous (Wi-Fi/LTE) uplink networks. It introduces, for the first time, multi-objective Bayesian optimization to this domain, formulating path scheduling as a black-box problem that jointly minimizes makespan and LTE usage. Without requiring modifications to the protocol internals or predefined scheduling policies, the approach employs probabilistic path selection to automatically explore the Pareto-optimal front, fully characterizing the latency–cost trade-off spectrum. Experimental results using Mininet-WiFi demonstrate that, under highly congested conditions, the proposed method reduces LTE consumption by up to 80% with only modest and controllable increases in latency, substantially outperforming fixed-policy schedulers.
该研究通过组织三个不同赛道,评估了语音增强、情感TTS及口音TTS系统的主观语音质量自动预测方法,吸引了18个团队参与并超越基线。
This work addresses the challenge that conventional time-invariant Koopman operators struggle to effectively model frequency-dependent dynamics in non-stationary time series. To overcome this limitation, the authors propose NDKoop, an end-to-end neural architecture that, for the first time, unifies learnable signal decomposition and Koopman-based dynamic modeling within a single framework. By explicitly separating the trend component (frequency-independent) from periodic components (frequency-dependent) and modeling each with dedicated Koopman networks, NDKoop achieves a more accurate representation of non-stationary dynamics. Extensive experiments demonstrate that NDKoop significantly outperforms existing methods across multiple time series forecasting benchmarks, confirming the efficacy and superiority of the proposed decomposition strategy in scenarios where ideal linearization assumptions do not hold.
This work addresses the limited generalization of existing model inversion attacks, which struggle to recover precise identity information from high-dimensional data. The authors propose GradLock, a training-time injection-based attack that stealthily embeds sensitive training samples into the AI supply chain and employs a dynamic gradient-locking mechanism to prevent payload degradation during optimization, enabling pixel-perfect reconstruction of original data without access to the training environment. GradLock achieves near-instantaneous extraction (<1.0 second), exhibits strong robustness against common deployment optimizations—including quantization, pruning, and fine-tuning—and leverages a stateless deterministic indexing scheme to construct an isolated data vault ensuring payload integrity. Experiments on MNIST, Imagenette, and CelebA demonstrate near-perfect reconstruction (SSIM ≈ 1.0), with 93.3% of users failing to detect the malicious logic, thereby exposing a critical security blind spot in the AI supply chain.
This study addresses foreign language anxiety (FLA), a significant barrier to second language acquisition—particularly in conversational contexts—exacerbated by dialogue agents that often produce overly complex utterances. To mitigate this, the authors propose a multi-agent embodied dialogue system that dynamically adapts linguistic complexity through an innovative “generate–evaluate–regenerate” loop, integrating CEFR proficiency levels with a level classifier to simplify output in real time. This approach uniquely combines multi-agent collaboration with CEFR-aligned simplification, enhancing both comprehensibility and psychological comfort. In a small-scale experiment with Japanese university students, 87.4% of generated utterances fell within ±1 CEFR level of learners’ self-assessed proficiency, substantially outperforming an unsimplified baseline (54.1%), thereby demonstrating the system’s efficacy in tailoring dialogue difficulty to individual learners.
This work addresses the challenge of balancing upload latency and cellular cost in MPQUIC scheduling over heterogeneous (Wi-Fi/LTE) uplink networks. It introduces, for the first time, multi-objective Bayesian optimization to this domain, formulating path scheduling as a black-box problem that jointly minimizes makespan and LTE usage. Without requiring modifications to the protocol internals or predefined scheduling policies, the approach employs probabilistic path selection to automatically explore the Pareto-optimal front, fully characterizing the latency–cost trade-off spectrum. Experimental results using Mininet-WiFi demonstrate that, under highly congested conditions, the proposed method reduces LTE consumption by up to 80% with only modest and controllable increases in latency, substantially outperforming fixed-policy schedulers.