Domain Fine-Tuning FinBERT on Finnish Histopathological Reports: Train-Time Signals and Downstream Correlations
This study addresses the scarcity of annotated Finnish medical texts by proposing a method to predict the benefit of unsupervised pretraining for downstream tasks through analyzing geometric changes in the embedding space during domain-adaptive fine-tuning. Specifically, we adapt FinBERT to Finnish pathology reports and investigate the relationship between the dynamics of embedding geometry and supervised classification performance. Our experiments demonstrate that the evolution of the embedding space in early fine-tuning stages effectively predicts the final model performance. This finding offers a practical approach for evaluating the efficacy of fine-tuning in medical AI settings—where acquiring labeled data is costly—without requiring annotated examples upfront.