Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of audio understanding in infant-centered naturalistic recordings under low signal-to-noise ratios, scarce annotations, and cross-household domain shifts. The authors propose a household-conditioned, multi-level audio tagger that integrates a LoRA-finetuned Whisper encoder with a lightweight speaker-aware Transformer. By introducing factorized speaker tokens and a sequence-level smoothing loss, the model enables frame-level predictions over long contextual segments. The structured speaker-conditional modeling effectively mitigates household-specific biases, substantially enhancing cross-household generalization and temporal consistency. Evaluated in real-world daytime home environments, the approach achieves efficient and accurate multi-level semantic annotation.
📝 Abstract
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.
Problem

Research questions and friction points this paper is trying to address.

infant-centered audio
low signal-to-noise ratio
limited labeled data
cross-family domain shifts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Whisper
LoRA
speaker conditioning
multi-tier audio tagging
domain generalization
🔎 Similar Papers