MIDAS: Mixing Ambiguous Data with Soft Labels for Dynamic Facial Expression Recognition
In dynamic facial expression recognition (DFER) under unconstrained real-world conditions, motion blur and semantic ambiguity of expressions severely hinder accurate classification. To address this, we propose the first video-level soft-label mixup augmentation method, which jointly performs convex interpolation across video frames and their corresponding multi-emotion probability soft labels—explicitly modeling both expression continuity and semantic uncertainty. Our approach comprises three components: (1) soft label construction via emotion distribution estimation, (2) soft-label-guided frame-level mixup augmentation, and (3) an end-to-end trainable framework. Evaluated on the DFEW benchmark, our method achieves significant improvements over existing state-of-the-art methods, demonstrating that soft-label mixing enhances model robustness to ambiguous, dynamically evolving expressions in the wild. This work establishes a novel paradigm for uncertainty-aware learning in DFER, advancing the integration of probabilistic semantics into video-based representation learning.