Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of rigorous quantification for emotion recognition accuracy ceilings and the difficulty in determining model saturation by proposing BACE, a Bias-Corrected Estimation framework. Integrating anchored Dirichlet mixture Bayesian estimation with noise deconvolution, this method introduces a fixed-declaration gating mechanism to prevent circular reasoning, thereby achieving the first separation of reducible and irreducible errors with calibrated confidence. Experiments on datasets such as GoEmotions demonstrate that at least 33% of classification errors are irreducible, revealing inherent limitations of single estimators. Consequently, this work provides a rigorous theoretical foundation and quantitative tooling for evaluating model performance bottlenecks in affective computing.
📝 Abstract
Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline. Our aim is not to pin this ceiling to a single number, but to quantify how far it depends on finite annotation, estimator choice, annotation noise, and the evaluation protocol, and thereby to discipline how confidently saturation can be claimed. We propose Bias-corrected Affective Ceiling Estimation (BACE), an analysis framework that estimates a bias-corrected ceiling, separates irreducible from reducible error, and disciplines the resulting claims. An anchored Dirichlet-mixture empirical Bayes estimator, bracketed between plug-in and NSB, recovers the human-consensus distribution; an annotator split, a noise deconvolution, and a fixed claim gate then attribute error without circularity. Methodologically, unconstrained point estimates place reachability anywhere from 0.38 to 1.03, so saturation cannot be decided by any single estimator. Substantively, the only assertion passing the claim gate is that at least about 33% of a representative classifier's error on GoEmotions is irreducible, with the same pattern recurring on offensiveness and irony.
Problem

Research questions and friction points this paper is trying to address.

Emotion Predictability Ceiling
Human Label Variation
Irreducible Error
Annotation Noise
Performance Saturation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bias-Corrected Ceiling Estimation
Instance-Level Fano Bounds
Dirichlet-Mixture Empirical Bayes
Noise Deconvolution
Irreducible Error