PRiSM: Benchmarking Phone Realization in Speech Models
Current evaluations of phoneme recognition systems are largely confined to surface-level transcription accuracy, with limited insight into their underlying phonemic perception capabilities. This work proposes PRiSM—the first open-source comprehensive benchmark for phoneme recognition—which establishes a standardized evaluation framework integrating both intrinsic (representation probing) and extrinsic (downstream tasks across clinical, educational, and multilingual settings) assessments. Leveraging transcription-based metrics, multilingual datasets, and encoder-CTC architectures, the benchmark enables reproducible evaluation and reveals that multilingual training substantially enhances performance, encoder-CTC models exhibit the most consistent results, and specialized phoneme recognition models still outperform general-purpose large audio language models.