Institution profile

Max Planck Institute for Evolutionary Anthropology

Academic institutioneurope · de
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Using Correspondence Patterns to Identify Irregular Words in Cognate sets Through Leave-One-Out Validation

Feb 02, 2026

This study addresses the lack of quantitative evaluation of phonological correspondence regularity in historical linguistic comparison, which hinders the effective identification of irregular word forms within cognate sets. The authors propose a novel regularity metric based on balanced mean recall and integrate it with leave-one-out validation and subsampling strategies to develop an automated method for detecting irregular forms. This work introduces balanced mean recall into computational historical linguistics for the first time, effectively capturing systematic deviations in sound correspondence patterns. Evaluated on real-world cognate datasets, the method achieves 85% accuracy in identifying irregular forms, substantially improving the data quality of cognate sets and providing a more reliable foundation for historical language reconstruction.

0 citationsRead paper

Most over-representation of phonological features in basic vocabulary disappears when controlling for spatial and phylogenetic effects

Dec 08, 2025

This study investigates whether sound symbolism patterns exhibit cross-linguistic universality robust to phylogenetic and areal dependencies. Using basic vocabulary from 2,864 languages in the Lexibank database, we apply phylogenetic independent contrasts and spatial autocorrelation modeling to rigorously disentangle genetic inheritance from areal contact confounds. Results show that most previously reported sound–meaning associations—such as /t/ with ‘pointed’ or /ŋ/ with ‘nose’—fail to survive stringent multilevel controls; only a few, including /m/ with ‘mother’ and /i/ with ‘small’, remain statistically robust. This is the first large-scale systematic test of sound symbolism universality, demonstrating that uncontrolled confounding factors critically inflate spurious cross-linguistic correlations. The findings underscore the necessity of hierarchical confound control in linguistic universals research and provide the strongest empirical constraints to date on sound symbolism theory.

0 citationsRead paper

Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data

Mar 14, 2025

Existing cross-linguistic colexification databases suffer from imbalanced language family coverage, absent phonetic transcription, and weak cross-linguistic comparability. To address these limitations, this study introduces a next-generation standardized colexification database. Methodologically, we propose a novel workflow integrating phylogenetically balanced sampling with full-scale IPA transcription, augmented by structured data modeling and versioned quality assessment. The resulting resource expands language family coverage by 42%, achieves 100% IPA standardization for all lexical forms, and substantially enhances cross-linguistic comparability and computational usability. For the first time, it unifies breadth—encompassing 32 global language families—with precision—providing fine-grained phonemic representations. This database has become a benchmark dataset across multiple disciplines, including linguistic typology, historical linguistics, psycholinguistics, and computational linguistics.

0 citationsRead paper

From Isolates to Families: Using Neural Networks for Automated Language Affiliation

Feb 17, 2025

In historical linguistics, the genetic classification of isolates and unclassified languages has long relied on labor-intensive manual comparison. This paper proposes the first multimodal deep neural network model integrating lexical and grammatical features, trained on morphological and syntactic data from over 1,000 languages worldwide to enable automatic cross-family classification. Methodologically, it jointly encodes lexical items (e.g., cognate sets) and structural properties (e.g., word order, alignment) within a unified architecture. Key contributions include: (1) the first systematic integration of lexical and grammatical representations, empirically confirming lexical features’ dominance in deep genetic inference; (2) interpretable attribution analyses that uncover plausible genealogical links for isolates; and (3) substantial gains over unimodal baselines—particularly in detecting distant subfamily relationships and assigning preliminary classifications to undocumented languages. The framework establishes a scalable, interpretable computational paradigm for deep phylogenetic reconstruction.

0 citationsRead paper
Recent publications

Latest Papers

Using Correspondence Patterns to Identify Irregular Words in Cognate sets Through Leave-One-Out Validation

Feb 02, 2026

This study addresses the lack of quantitative evaluation of phonological correspondence regularity in historical linguistic comparison, which hinders the effective identification of irregular word forms within cognate sets. The authors propose a novel regularity metric based on balanced mean recall and integrate it with leave-one-out validation and subsampling strategies to develop an automated method for detecting irregular forms. This work introduces balanced mean recall into computational historical linguistics for the first time, effectively capturing systematic deviations in sound correspondence patterns. Evaluated on real-world cognate datasets, the method achieves 85% accuracy in identifying irregular forms, substantially improving the data quality of cognate sets and providing a more reliable foundation for historical language reconstruction.

0 citationsRead paper

Most over-representation of phonological features in basic vocabulary disappears when controlling for spatial and phylogenetic effects

Dec 08, 2025

This study investigates whether sound symbolism patterns exhibit cross-linguistic universality robust to phylogenetic and areal dependencies. Using basic vocabulary from 2,864 languages in the Lexibank database, we apply phylogenetic independent contrasts and spatial autocorrelation modeling to rigorously disentangle genetic inheritance from areal contact confounds. Results show that most previously reported sound–meaning associations—such as /t/ with ‘pointed’ or /ŋ/ with ‘nose’—fail to survive stringent multilevel controls; only a few, including /m/ with ‘mother’ and /i/ with ‘small’, remain statistically robust. This is the first large-scale systematic test of sound symbolism universality, demonstrating that uncontrolled confounding factors critically inflate spurious cross-linguistic correlations. The findings underscore the necessity of hierarchical confound control in linguistic universals research and provide the strongest empirical constraints to date on sound symbolism theory.

0 citationsRead paper

Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data

Mar 14, 2025

Existing cross-linguistic colexification databases suffer from imbalanced language family coverage, absent phonetic transcription, and weak cross-linguistic comparability. To address these limitations, this study introduces a next-generation standardized colexification database. Methodologically, we propose a novel workflow integrating phylogenetically balanced sampling with full-scale IPA transcription, augmented by structured data modeling and versioned quality assessment. The resulting resource expands language family coverage by 42%, achieves 100% IPA standardization for all lexical forms, and substantially enhances cross-linguistic comparability and computational usability. For the first time, it unifies breadth—encompassing 32 global language families—with precision—providing fine-grained phonemic representations. This database has become a benchmark dataset across multiple disciplines, including linguistic typology, historical linguistics, psycholinguistics, and computational linguistics.

0 citationsRead paper

From Isolates to Families: Using Neural Networks for Automated Language Affiliation

Feb 17, 2025

In historical linguistics, the genetic classification of isolates and unclassified languages has long relied on labor-intensive manual comparison. This paper proposes the first multimodal deep neural network model integrating lexical and grammatical features, trained on morphological and syntactic data from over 1,000 languages worldwide to enable automatic cross-family classification. Methodologically, it jointly encodes lexical items (e.g., cognate sets) and structural properties (e.g., word order, alignment) within a unified architecture. Key contributions include: (1) the first systematic integration of lexical and grammatical representations, empirically confirming lexical features’ dominance in deep genetic inference; (2) interpretable attribution analyses that uncover plausible genealogical links for isolates; and (3) substantial gains over unimodal baselines—particularly in detecting distant subfamily relationships and assigning preliminary classifications to undocumented languages. The framework establishes a scalable, interpretable computational paradigm for deep phylogenetic reconstruction.

0 citationsRead paper