🤖 AI Summary
This study investigates whether large language models (LLMs), trained solely on surface-level sequences, can spontaneously acquire hierarchical syntactic sensitivity—specifically, whether they exhibit structure-dependent generalization without explicit grammatical supervision. Method: Focusing on two canonical structure-dependent phenomena—subject–auxiliary inversion and parasitic gap licensing—the authors employ controlled prompt engineering to elicit grammaticality judgments from models including GPT-4 and LLaMA-3, then rigorously assess structural generalization via systematic minimal-pair comparisons. Contribution/Results: Results demonstrate that LLMs consistently distinguish grammatical from ungrammatical sentences across novel constructions, exhibiting functional syntactic competence that transcends linear sequential cues. Critically, this hierarchical structural sensitivity emerges robustly despite the absence of overt syntactic annotations or architectural biases toward hierarchy. This work provides the first systematic empirical evidence that LLMs spontaneously develop hierarchical syntactic representations from surface input alone—challenging the generative grammar tenet that syntactic structure must be innately encoded.
📝 Abstract
What counts as evidence for syntactic structure? In traditional generative grammar, systematic contrasts in grammaticality such as subject-auxiliary inversion and the licensing of parasitic gaps are taken as evidence for an internal, hierarchical grammar. In this paper, we test whether large language models (LLMs), trained only on surface forms, reproduce these contrasts in ways that imply an underlying structural representation.
We focus on two classic constructions: subject-auxiliary inversion (testing recognition of the subject boundary) and parasitic gap licensing (testing abstract dependency structure). We evaluate models including GPT-4 and LLaMA-3 using prompts eliciting acceptability ratings. Results show that LLMs reliably distinguish between grammatical and ungrammatical variants in both constructions, and as such support that they are sensitive to structure and not just linear order. Structural generalizations, distinct from cognitive knowledge, emerge from predictive training on surface forms, suggesting functional sensitivity to syntax without explicit encoding.