Zero-Shot Grammar Competency Estimation Using Large Language Model Generated Pseudo Labels
Expert annotations for grammatical proficiency assessment are scarce—especially for spontaneous, disfluent spoken language—hindering reliable evaluation. Method: This paper proposes a zero-shot pseudo-labeling framework that leverages scoring rubrics to guide large language models (LLMs) in generating high-quality pseudo-labels. It integrates a noise-robust training mechanism and a score-consistency optimization strategy within a Transformer-based architecture to ensure robust modeling. Contribution/Results: To our knowledge, this is the first work to jointly employ structured prompt-driven pseudo-label generation and noise-robust learning for grammatical proficiency estimation—requiring no human annotations and supporting both spoken and written modalities. Experiments demonstrate substantial improvements over supervised baselines in low-resource settings, achieving high accuracy, strong interpretability (via rubric-aligned outputs), and practical applicability for real-world language assessment.