๐ค AI Summary
This study evaluates the biological fidelity of the self-supervised audiovisual speech model AV-HuBERT in modeling human multisensory integration, particularly the McGurk effect. Using a standard McGurk experimental paradigm and statistical analyses, the authors compare AV-HuBERTโs responses to incongruent audiovisual stimuli against behavioral data from 44 human participants. The model exhibits a striking alignment with human perception in auditory dominance rates (32.0% vs. 31.8%) but demonstrates significantly higher phoneme fusion rates (68.0% vs. 47.7%). Moreover, it lacks the variability and diversity of perceptual errors characteristic of human responses. These findings indicate that while AV-HuBERT captures certain aspects of human multisensory speech integration, it remains limited by rigid decision-making mechanisms that fail to replicate the flexibility and stochasticity inherent in human perception.
๐ Abstract
This study evaluates AV-HuBERT's perceptual bio-fidelity by benchmarking its response to incongruent audiovisual stimuli (McGurk effect) against human observers (N=44). Results reveal a striking quantitative isomorphism: AI and humans exhibited nearly identical auditory dominance rates (32.0% vs. 31.8%), suggesting the model captures biological thresholds for auditory resistance. However, AV-HuBERT showed a deterministic bias toward phonetic fusion (68.0%), significantly exceeding human rates (47.7%). While humans displayed perceptual stochasticity and diverse error profiles, the model remained strictly categorical. Findings suggest that current self-supervised architectures mimic multisensory outcomes but lack the neural variability inherent to human speech perception.