Speech transformer models for extracting information from baby cries

📅 2025-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Infant cry analysis faces challenges including acoustic instability, data scarcity, and infant identity disambiguation—tasks poorly addressed by conventional speech models trained exclusively on linguistic signals. Method: This work systematically investigates the transferability and representational properties of Transformer-based pretrained speech models on non-speech infant cry signals. We evaluate multiple models across eight diverse datasets comprising 115 hours of audio from 960 infants, targeting cry classification, vocalization characteristic modeling, and infant identity recognition. Contribution/Results: We demonstrate that pretrained speech representations effectively encode physiological state and speaker-identity information in cries. Model architecture and pretraining strategy critically influence cross-domain generalization. Crucially, our findings reveal that speech self-supervised models implicitly learn acoustic–physiological mappings—capturing biologically grounded structure beyond phonetic content. This provides an interpretable representation foundation and a novel model design paradigm for affective computing and early-life health monitoring.

Technology Category

Application Category

📝 Abstract
Transfer learning using latent representations from pre-trained speech models achieves outstanding performance in tasks where labeled data is scarce. However, their applicability to non-speech data and the specific acoustic properties encoded in these representations remain largely unexplored. In this study, we investigate both aspects. We evaluate five pre-trained speech models on eight baby cries datasets, encompassing 115 hours of audio from 960 babies. For each dataset, we assess the latent representations of each model across all available classification tasks. Our results demonstrate that the latent representations of these models can effectively classify human baby cries and encode key information related to vocal source instability and identity of the crying baby. In addition, a comparison of the architectures and training strategies of these models offers valuable insights for the design of future models tailored to similar tasks, such as emotion detection.
Problem

Research questions and friction points this paper is trying to address.

Evaluating pre-trained speech models on baby cry classification tasks
Assessing latent representations for encoding acoustic properties in cries
Exploring model applicability to non-speech vocalization data analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transfer learning from pre-trained speech models
Evaluating latent representations on baby cry datasets
Encoding vocal instability and identity information
💼 Related Jobs
No related jobs found.
G
Guillem Bonafos
ENES Bioacoustics Research Lab
J
Jéremy Rouch
ENES Bioacoustics Research Lab
L
Lény Lego
ENES Bioacoustics Research Lab
D
David Reby
ENES Bioacoustics Research Lab and Institut Universitaire de France
H
Hugues Patural
SAINBIOSE Laboratory
N
Nicolas Mathevon
ENES Bioacoustics Research Lab, CHArt Lab, Institut Universitaire de France
R
Rémy Emonet
Laboratoire Hubert Curien, Institut Universitaire de France