I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用GPT-5模型对幼儿课堂师生互动进行评分,并与人工评分比较,发现AI在情感支持领域表现较好,但在组织和教学支持方面差异较大。
📝 Abstract
Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study evaluated the feasibility of using a LLM (GPT-5 model) to score teacher-child interactions in early childhood classrooms, benchmarked against human raters. The study analyzed 87 video-recorded observations from 38 classrooms across 30 kindergartens in Hong Kong. Using observation transcripts, the AI model was configured to apply the full Classroom Assessment Scoring System (CLASS) framework. AI-rated scores were then compared with human ratings by examining correlations and differences in mean scores of the CLASS domains and dimensions. The results showed greater convergence between AI and raters for the Emotional Support domain and, in particular, the Quality of Feedback dimension, which captures how teachers use feedback to extend children's learning. Greater divergence emerged for interactions that were more procedural or context-dependent, particularly within the Classroom Organization and Instructional Support domains. These findings suggest that transcript-based AI scoring may capture some of the relative variation in teacher-child interactions but cannot yet reproduce calibrated human judgements consistently across the full CLASS framework. AI-assisted observation may therefore be more appropriate as a preliminary screening tool rather than as a replacement for trained observers, providing teachers with evidence for reflection rather than high-stakes evaluation. Future research should examine whether domain-specific training and incorporation of contextual and visual information can improve alignment between AI and human rated scores.
Problem

Research questions and friction points this paper is trying to address.

classroom observations
AI-rated scores
teacher-child interactions
CLASS framework
human raters
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM
GPT-5
CLASS framework
Emotional Support
Quality of Feedback
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yasmin Fong
Department of Early Childhood Education, The Education University of Hong Kong, 10 Lo Ping Road, Tai Po, Hong Kong SAR, China
J
Jane Xiang
Department of Early Childhood Education, The Education University of Hong Kong, 10 Lo Ping Road, Tai Po, Hong Kong SAR, China
T
Tak-Yue Dickson Chan
Department of Early Childhood Education, The Education University of Hong Kong, 10 Lo Ping Road, Tai Po, Hong Kong SAR, China
K
Kerry Lee
Yew Chung College of Early Childhood Education, 2 Tin Wan Hill Road, Tin Wan, Hong Kong SAR, China
E
Eva Yi Hung Lau
Department of Early Childhood Education, The Education University of Hong Kong, 10 Lo Ping Road, Tai Po, Hong Kong SAR, China