Mi\'{c}i Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect
This study addresses the scarcity of structured speech–text aligned data for the endangered Chakavian dialect, which has hindered its integration into artificial intelligence applications. We present the first high-quality, word-aligned multimodal dataset of *The Little Prince* in Chakavian, comprising synchronized text, images, and audio, meticulously curated through manual alignment and publicly released via the CLARIN.SI platform. Fine-tuning the Whisper-large-v3 model on this dataset yields substantial improvements in automatic speech recognition performance, reducing the word error rate by 50% and the character error rate by approximately two-thirds on the test set. This work establishes a reproducible data paradigm and technical framework for AI-driven preservation of endangered dialects.