Exploring Second-Order Pattern Recognition in Speaker Recognition

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过应用层次聚类算法和HCNA方法,探索了语音识别中第二阶模式的发现与识别问题。
📝 Abstract
In classical pattern recognition tasks, neural networks are trained to recognise human-defined patterns for model inputs. Some Explainable AI (XAI) methods can explain other latent patterns that underlie the network's recognition of inputs as human-defined patterns; in this work, we call these latent patterns second-order patterns, and we propose to discover them. To this end, we apply a hierarchical clustering algorithm to analyse whether representations learned by a speaker recognition network from utterances naturally form hierarchical clusters. Each resulting cluster represents a second-order pattern that characterises how the network recognises some known utterances as speaker identities. All the resulting second-order patterns are then semantically interpreted using the existing Hierarchical Cluster-Class Matching (HCCM) method. Furthermore, we propose a new task, second-order pattern recognition, to identify which discovered second-order patterns characterising known utterances are exhibited by an unseen utterance. To achieve this, we design the Hierarchical Cluster Navigation and Assignment (HCNA) method. HCNA recognises a known second-order pattern as applying to an unseen utterance when the unseen utterance's network representation lies within the extrapolation space of the cluster regarded as that second-order pattern. Our experiments show that the extrapolation mechanism introduced by HCNA substantially improves performance on the second-order pattern recognition task.
Problem

Research questions and friction points this paper is trying to address.

second-order pattern
speaker recognition
hierarchical clustering
Innovation

Methods, ideas, or system contributions that make the work stand out.

second-order pattern recognition
Hierarchical Cluster Navigation and Assignment (HCNA)
extrapolation mechanism
🔎 Similar Papers
No similar papers found.
Y
Yanze Xu
Centre for Vision, Speech and Signal Processing, University of Surrey, Guildford, UK
Wenwu Wang
Wenwu Wang
Professor, University of Surrey, UK
signal processingmachine learningmachine listeningaudio/speech/audio-visualmultimodal fusion
M
Mark D. Plumbley
Department of Informatics, King’s College London, London, UK