Integrated ensemble of BERT- and features-based models for authorship attribution in Japanese literary works
This paper addresses the few-shot Japanese literary author attribution (AA) task. Methodologically, it proposes a dual-path ensemble framework integrating traditional stylistic features with pre-trained language models (PLMs). It is the first to empirically validate BERT’s effectiveness for few-shot Japanese AA; simultaneously, TF-IDF and character n-gram features are extracted and modeled using multi-layer classifiers—XGBoost, SVM, and MLP—with weighted voting for fusion. The core contribution lies in establishing a synergistic “PLM + traditional features” dual-path paradigm, substantially improving generalization under data-scarce conditions. Experiments on test sets disjoint from pre-training data demonstrate an approximately 14-percentage-point improvement in macro-F1 score over baseline models. The integrated model consistently outperforms all individual components, establishing a new state-of-the-art performance ceiling for few-shot Japanese AA.