Contextual Embedding Evidence for Main--Light Verb Distinctions in Urdu

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用UrduBERT等模型的上下文嵌入来区分乌尔都语中的主要动词和轻动词,证明两者在表示上存在显著差异但保持词汇关联。
📝 Abstract
Urdu light verbs contribute schematic event-structural meaning while remaining lexically related to corresponding main verbs. This study tests representational predictions derived from Butt's analysis using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT across 1,126 naturally occurring sentences containing seven Urdu verbs. Main and light uses show significant representational separation in all 21 verb--model comparisons. At the same time, same-lemma main and light centroids are consistently closer than mismatched main--light lemma pairs, supporting continued lexical relatedness. In a seven-way prediction task restricted to light uses, verb identity remains recoverable after the target is masked, with UrduBERT achieving 0.866 accuracy and 0.852 macro-F1. UrduBERT also retains 0.782 accuracy under a preceding-form-disjoint evaluation, indicating generalization beyond repeated local verb combinations. These findings provide computational evidence consistent with Butt's account that Urdu light verbs differ systematically from their main uses while retaining lemma-specific and verb-specific representational structure.
Problem

Research questions and friction points this paper is trying to address.

Urdu light verbs
contextual embeddings
lexical relatedness
representational structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

contextual embeddings
Urdu light verbs
representational separation
lexical relatedness
generalization
🔎 Similar Papers
No similar papers found.