Multilevel modelling of double-sampled clustered social networks with individual-level data on between-cluster ties
本文提出了一种多层社会关系模型,通过双采样方法调整报告者效应,并利用MCMC方法估计模型参数,以解决不同集群间个体联系的网络分析问题。
本文提出了一种多层社会关系模型,通过双采样方法调整报告者效应,并利用MCMC方法估计模型参数,以解决不同集群间个体联系的网络分析问题。
This study addresses the lack of quantitative evaluation of phonological correspondence regularity in historical linguistic comparison, which hinders the effective identification of irregular word forms within cognate sets. The authors propose a novel regularity metric based on balanced mean recall and integrate it with leave-one-out validation and subsampling strategies to develop an automated method for detecting irregular forms. This work introduces balanced mean recall into computational historical linguistics for the first time, effectively capturing systematic deviations in sound correspondence patterns. Evaluated on real-world cognate datasets, the method achieves 85% accuracy in identifying irregular forms, substantially improving the data quality of cognate sets and providing a more reliable foundation for historical language reconstruction.
This study investigates whether sound symbolism patterns exhibit cross-linguistic universality robust to phylogenetic and areal dependencies. Using basic vocabulary from 2,864 languages in the Lexibank database, we apply phylogenetic independent contrasts and spatial autocorrelation modeling to rigorously disentangle genetic inheritance from areal contact confounds. Results show that most previously reported sound–meaning associations—such as /t/ with ‘pointed’ or /ŋ/ with ‘nose’—fail to survive stringent multilevel controls; only a few, including /m/ with ‘mother’ and /i/ with ‘small’, remain statistically robust. This is the first large-scale systematic test of sound symbolism universality, demonstrating that uncontrolled confounding factors critically inflate spurious cross-linguistic correlations. The findings underscore the necessity of hierarchical confound control in linguistic universals research and provide the strongest empirical constraints to date on sound symbolism theory.
Existing cross-linguistic colexification databases suffer from imbalanced language family coverage, absent phonetic transcription, and weak cross-linguistic comparability. To address these limitations, this study introduces a next-generation standardized colexification database. Methodologically, we propose a novel workflow integrating phylogenetically balanced sampling with full-scale IPA transcription, augmented by structured data modeling and versioned quality assessment. The resulting resource expands language family coverage by 42%, achieves 100% IPA standardization for all lexical forms, and substantially enhances cross-linguistic comparability and computational usability. For the first time, it unifies breadth—encompassing 32 global language families—with precision—providing fine-grained phonemic representations. This database has become a benchmark dataset across multiple disciplines, including linguistic typology, historical linguistics, psycholinguistics, and computational linguistics.
In historical linguistics, the genetic classification of isolates and unclassified languages has long relied on labor-intensive manual comparison. This paper proposes the first multimodal deep neural network model integrating lexical and grammatical features, trained on morphological and syntactic data from over 1,000 languages worldwide to enable automatic cross-family classification. Methodologically, it jointly encodes lexical items (e.g., cognate sets) and structural properties (e.g., word order, alignment) within a unified architecture. Key contributions include: (1) the first systematic integration of lexical and grammatical representations, empirically confirming lexical features’ dominance in deep genetic inference; (2) interpretable attribution analyses that uncover plausible genealogical links for isolates; and (3) substantial gains over unimodal baselines—particularly in detecting distant subfamily relationships and assigning preliminary classifications to undocumented languages. The framework establishes a scalable, interpretable computational paradigm for deep phylogenetic reconstruction.
本文提出了一种多层社会关系模型,通过双采样方法调整报告者效应,并利用MCMC方法估计模型参数,以解决不同集群间个体联系的网络分析问题。
This study addresses the lack of quantitative evaluation of phonological correspondence regularity in historical linguistic comparison, which hinders the effective identification of irregular word forms within cognate sets. The authors propose a novel regularity metric based on balanced mean recall and integrate it with leave-one-out validation and subsampling strategies to develop an automated method for detecting irregular forms. This work introduces balanced mean recall into computational historical linguistics for the first time, effectively capturing systematic deviations in sound correspondence patterns. Evaluated on real-world cognate datasets, the method achieves 85% accuracy in identifying irregular forms, substantially improving the data quality of cognate sets and providing a more reliable foundation for historical language reconstruction.
This study investigates whether sound symbolism patterns exhibit cross-linguistic universality robust to phylogenetic and areal dependencies. Using basic vocabulary from 2,864 languages in the Lexibank database, we apply phylogenetic independent contrasts and spatial autocorrelation modeling to rigorously disentangle genetic inheritance from areal contact confounds. Results show that most previously reported sound–meaning associations—such as /t/ with ‘pointed’ or /ŋ/ with ‘nose’—fail to survive stringent multilevel controls; only a few, including /m/ with ‘mother’ and /i/ with ‘small’, remain statistically robust. This is the first large-scale systematic test of sound symbolism universality, demonstrating that uncontrolled confounding factors critically inflate spurious cross-linguistic correlations. The findings underscore the necessity of hierarchical confound control in linguistic universals research and provide the strongest empirical constraints to date on sound symbolism theory.
Existing cross-linguistic colexification databases suffer from imbalanced language family coverage, absent phonetic transcription, and weak cross-linguistic comparability. To address these limitations, this study introduces a next-generation standardized colexification database. Methodologically, we propose a novel workflow integrating phylogenetically balanced sampling with full-scale IPA transcription, augmented by structured data modeling and versioned quality assessment. The resulting resource expands language family coverage by 42%, achieves 100% IPA standardization for all lexical forms, and substantially enhances cross-linguistic comparability and computational usability. For the first time, it unifies breadth—encompassing 32 global language families—with precision—providing fine-grained phonemic representations. This database has become a benchmark dataset across multiple disciplines, including linguistic typology, historical linguistics, psycholinguistics, and computational linguistics.
In historical linguistics, the genetic classification of isolates and unclassified languages has long relied on labor-intensive manual comparison. This paper proposes the first multimodal deep neural network model integrating lexical and grammatical features, trained on morphological and syntactic data from over 1,000 languages worldwide to enable automatic cross-family classification. Methodologically, it jointly encodes lexical items (e.g., cognate sets) and structural properties (e.g., word order, alignment) within a unified architecture. Key contributions include: (1) the first systematic integration of lexical and grammatical representations, empirically confirming lexical features’ dominance in deep genetic inference; (2) interpretable attribution analyses that uncover plausible genealogical links for isolates; and (3) substantial gains over unimodal baselines—particularly in detecting distant subfamily relationships and assigning preliminary classifications to undocumented languages. The framework establishes a scalable, interpretable computational paradigm for deep phylogenetic reconstruction.