LLM Agents as Computational Typologists

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入基于LLM的AUTOTYPOLOGIST系统,解决了大规模跨语言比较劳动密集且难以扩展的问题,该系统能自动进行语法分析和类型学假设检验。
📝 Abstract
Linguistic typology relies on expert analysis of reference grammars across languages, making large-scale crosslinguistic comparison labor-intensive and unscalable. We introduce AUTOTYPOLOGIST, an LLM agent for evidence-grounded typological analysis over reference grammars. The agent is capable of retrieving relevant grammar sections, analyzing interlinear glossed text (IGT), and iteratively reasoning over typological hypotheses using a ReAct-style workflow. We evaluate the system on TYPOLOGICAL FEATURE CODING against expert annotations and TYPOLOGICAL HYPOTHESIS TESTING with typological universals using 25 open-source reference grammars. Operating under different information constraints in TYPOLOGICAL FEATURE CODING, the agent can synthesize information from reference grammar prose but still faces challenges with only IGTs in the target language. In TYPOLOGICAL HYPOTHESIS TESTING, the agent can synthesize crosslinguistic evidence and identify both supporting cases and counterexamples. These findings suggest that LLM agents can support scalable and inspectable typological analysis, while still requiring expert validation.
Problem

Research questions and friction points this paper is trying to address.

Linguistic Typology
Reference Grammars
Crosslinguistic Comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agent
ReAct-style Workflow
Typological Hypothesis Testing
Crosslinguistic Comparison
🔎 Similar Papers