Dec 06, 2025
Existing facial expression recognition (FER) datasets predominantly provide only discrete emotion category labels, failing to capture fine-grained, continuous affective variations. Although some works incorporate Valence-Arousal (VA) annotations, the Dominance (D) dimension remains largely absent. This paper addresses this gap by introducing the first complete, human-verified three-dimensional Valence-Arousal-Dominance (VAD) continuous annotation for the FER2013 dataset. We further propose an orthogonal convolution-based ResNet regression architecture that enforces feature orthogonality to improve VAD value prediction accuracy. Experiments demonstrate that, despite its high annotation difficulty, the D dimension is effectively learnable; orthogonal convolutions significantly enhance predictive performance across all three dimensions—particularly for Dominance. The released VAD-annotated FER2013 dataset and open-source code establish a new benchmark for multidimensional affective computing, enabling more precise, granular emotion analysis in applications such as intelligent education.