🤖 AI Summary
This work addresses the challenges of audio understanding in infant-centered naturalistic recordings under low signal-to-noise ratios, scarce annotations, and cross-household domain shifts. The authors propose a household-conditioned, multi-level audio tagger that integrates a LoRA-finetuned Whisper encoder with a lightweight speaker-aware Transformer. By introducing factorized speaker tokens and a sequence-level smoothing loss, the model enables frame-level predictions over long contextual segments. The structured speaker-conditional modeling effectively mitigates household-specific biases, substantially enhancing cross-household generalization and temporal consistency. Evaluated in real-world daytime home environments, the approach achieves efficient and accurate multi-level semantic annotation.
📝 Abstract
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.