🤖 AI Summary
This study addresses the poor performance of existing sentiment polarity models on structurally complex and highly heterogeneous large-scale oral history texts—such as Holocaust testimonies—under domain shift. The authors propose an ABC hierarchical framework grounded in inter-model consistency, integrating multi-model label triangulation with emotion distribution–based auxiliary analysis. They systematically evaluate 107,305 utterances and 579,013 sentences using three pretrained Transformer-based polarity classifiers alongside a T5 emotion classifier, quantifying model disagreement through agreement rates, Cohen’s Kappa coefficients, and normalized confusion matrices. Results reveal only low-to-moderate overall model consistency, with disagreements predominantly concentrated near the neutral boundary. Crucially, emotion distributions vary significantly across ABC hierarchy levels, offering an actionable diagnostic pathway for assessing uncertainty in affective analysis of sensitive historical narratives.
📝 Abstract
Polarity detection becomes substantially more challenging under domain shift, particularly in heterogeneous, long-form narratives with complex discourse structure, such as Holocaust oral histories. This paper presents a corpus-scale diagnostic study of off-the-shelf sentiment classifiers on long-form Holocaust oral histories, using three pretrained transformer-based polarity classifiers on a corpus of 107,305 utterances and 579,013 sentences. After assembling model outputs, we introduce an agreement-based stability taxonomy (ABC) to stratify inter-model output stability. We report pairwise percent agreement, Cohen kappa, Fleiss kappa, and row-normalized confusion matrices to localize systematic disagreement. As an auxiliary descriptive signal, a T5-based emotion classifier is applied to stratified samples from each agreement stratum to compare emotion distributions across strata. The combination of multi-model label triangulation and the ABC taxonomy provides a cautious, operational framework for characterizing where and how sentiment models diverge in sensitive historical narratives. Inter-model agreement is low to moderate overall and is driven primarily by boundary decisions around neutrality.