Entropy of Ukrainian
This study addresses a gap in information-theoretic research on Ukrainian by applying Shannon’s 1951 human prediction experiment methodology to this underexplored language. Recruiting 184 volunteers via crowdsourcing, the authors conducted an online character-level prediction task to estimate an upper bound on the language’s entropy. The experiment yielded an entropy upper bound of approximately 1.201 bits per character. The work presents the first human-prediction-based estimate of information entropy for Ukrainian and provides a fully reproducible framework, including open-sourced experimental protocols, code, and detailed challenge analyses. Furthermore, the results enable direct comparison with large language model performance, establishing a replicable paradigm for information-theoretic investigations of low-resource languages.