Institution profile

University of Belgrade

Academic institutioneurope · rs
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

Wiki Dumps to Training Corpora: South Slavic Case

Apr 28, 2026

This work proposes a systematic methodology for constructing high-quality training corpora for South Slavic language models from raw Wikimedia data. Starting with multilingual Wiki project texts, the approach first parses Wiki markup to extract natural language content and then employs an n-gram–based redundancy detection mechanism to effectively filter out highly repetitive, low-information articles. This pipeline substantially enhances the linguistic richness and authenticity of the resulting corpus while maintaining cross-lingual applicability. The final resource encompasses seven South Slavic languages, offering a reliable foundation for large language model training and cross-linguistic comparative studies.

0 citationsRead paper

Testing independence in the presence of missing data: high-dimensional case

Apr 24, 2026

This study addresses the challenge of testing variable independence in high-dimensional data with missing values by extending the nonparametric Kendall’s rank correlation framework to settings involving incomplete observations. The authors propose two novel corrected test statistics specifically designed to accommodate missingness, effectively integrating high-dimensional inference with explicit modeling of the missing data mechanism. Theoretical analysis establishes the statistical validity of the proposed methods, while extensive simulations demonstrate their robustness and high power across various missingness mechanisms—including both missing at random and not missing at random scenarios. This work substantially enhances the reliability and applicability of independence testing in high-dimensional settings where data incompleteness is prevalent.

0 citationsRead paper
Recent publications

Latest Papers

Wiki Dumps to Training Corpora: South Slavic Case

Apr 28, 2026

This work proposes a systematic methodology for constructing high-quality training corpora for South Slavic language models from raw Wikimedia data. Starting with multilingual Wiki project texts, the approach first parses Wiki markup to extract natural language content and then employs an n-gram–based redundancy detection mechanism to effectively filter out highly repetitive, low-information articles. This pipeline substantially enhances the linguistic richness and authenticity of the resulting corpus while maintaining cross-lingual applicability. The final resource encompasses seven South Slavic languages, offering a reliable foundation for large language model training and cross-linguistic comparative studies.

0 citationsRead paper

Testing independence in the presence of missing data: high-dimensional case

Apr 24, 2026

This study addresses the challenge of testing variable independence in high-dimensional data with missing values by extending the nonparametric Kendall’s rank correlation framework to settings involving incomplete observations. The authors propose two novel corrected test statistics specifically designed to accommodate missingness, effectively integrating high-dimensional inference with explicit modeling of the missing data mechanism. Theoretical analysis establishes the statistical validity of the proposed methods, while extensive simulations demonstrate their robustness and high power across various missingness mechanisms—including both missing at random and not missing at random scenarios. This work substantially enhances the reliability and applicability of independence testing in high-dimensional settings where data incompleteness is prevalent.

0 citationsRead paper