(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用跨模态泛化方法,通过视觉线索教授(V)LMs新名词,证明其能超越表面共现学习抽象规则。
📝 Abstract
Language models learn about grammatical number primarily from co-occurrence, and show frequency effects as a result---sometimes taken to indicate that they do not learn abstract ``rules'', and are instead dependent on specific lexical items. Testing generalization with text stimuli alone cannot settle this debate, since distributional cues (is/are, this/these) easily give number away. We instead use cross-modal generalization as a tool to investigate abstractions in LMs that can also accept visual inputs (VLMs), restricting the evidence that diagnoses number to an extra-linguistic modality. We teach VLMs pairs of new nouns by adding new embeddings and only updating them during learning, comparing conditions where number is diagnosed by visual cues alone against ones where it is disambiguated by text. Across behavior, representational dynamics, and causal mechanisms, we find non-trivial evidence for cross-modal generalization across both exposure conditions, and that linguistic vs. extra-linguistic cue conditions are treated in similar ways in the internal mechanisms of the model. This suggests that statistical learners like VLMs can generalize beyond surface-level co-occurrence and show genuine abstraction-compatible behavior.
Problem

Research questions and friction points this paper is trying to address.

cross-modal generalization
abstract rules
grammatical number
visual inputs
statistical learners
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal generalization
visual cues
linguistic abstraction
VLMs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.