Large Language Models for Summarizing Czech Historical Documents and Beyond
Historical document summarization for low-resource Czech has long been hindered by linguistic complexity and scarcity of annotated data. To address this, we introduce *Posel od Čerchova*, the first annotated summarization dataset for historical Czech, and achieve state-of-the-art performance on the modern Czech benchmark SumeCzech. Our approach integrates transfer learning with multilingual pretraining, adapting large language models—including Mistral and mT5—to jointly handle both modern and historical Czech texts, enabling the first unified cross-era summarization framework for the language. Key contributions are: (1) releasing the first open-source summarization dataset for historical Czech; (2) establishing the current strongest baseline for Czech summarization; and (3) empirically validating the efficacy of large language models in low-resource historical text NLP tasks, thereby providing a reusable methodological paradigm for analogous under-resourced languages.