Direct Preference Density Alignment for Conversational Audio Equalization
提出直接偏好密度对齐方法,利用大规模用户数据构建非参数偏好密度图,结合在线和离线优化优势,解决对话音频均衡问题。
提出直接偏好密度对齐方法,利用大规模用户数据构建非参数偏好密度图,结合在线和离线优化优势,解决对话音频均衡问题。
This work addresses the demand for low-latency, low-complexity single-channel speech enhancement on resource-constrained embedded devices by proposing an improved architecture. Specifically, the original GRU in ULCNet is replaced with a lightweight FastGRNN, and a novel trainable complementary filter is introduced to mitigate state drift during long-duration audio inference. The proposed method achieves speech enhancement performance comparable to the original ULCNet while reducing model size by over 50% and decreasing average inference latency by 34%. These improvements significantly enhance deployment efficiency and practical applicability on edge hardware without compromising perceptual quality or intelligibility.
提出直接偏好密度对齐方法,利用大规模用户数据构建非参数偏好密度图,结合在线和离线优化优势,解决对话音频均衡问题。
This work addresses the demand for low-latency, low-complexity single-channel speech enhancement on resource-constrained embedded devices by proposing an improved architecture. Specifically, the original GRU in ULCNet is replaced with a lightweight FastGRNN, and a novel trainable complementary filter is introduced to mitigate state drift during long-duration audio inference. The proposed method achieves speech enhancement performance comparable to the original ULCNet while reducing model size by over 50% and decreasing average inference latency by 34%. These improvements significantly enhance deployment efficiency and practical applicability on edge hardware without compromising perceptual quality or intelligibility.