Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of time-intensive and expertise-dependent human scoring in open-ended music analysis assessments by systematically evaluating the feasibility of automated scoring using a large language model (GPT-4o-mini). Leveraging 300 undergraduate responses and teacher-assigned benchmark scores, the authors compare three prompting strategies—few-shot with chain-of-thought (Fs+CoT), retrieval-augmented generation (RAG), and self-consistency (SC)—across four scoring dimensions through multiple experimental rounds to assess stability. Results indicate that Fs+CoT achieves the highest agreement with human raters both in single-run and median-aggregated settings; RAG exhibits a systematic tendency to overrate; SC demonstrates strong repeatability but weaker individual-level consistency; and scoring agreement is consistently lower on terminology than on reasoning dimensions. The findings underscore the critical role of prompting strategies in scoring validity and offer empirical support for automating educational assessment.
📝 Abstract
Scoring open-ended music analysis responses is time-consuming and requires nuanced judgments of harmonic knowledge and formal understanding. This study evaluates the validity and repeatability of GPT-4o-mini for rubric-based scoring of music analysis essays, using teacher mean scores as the benchmark. A dataset of 300 university-level student responses was scored by teachers on four dimensions: Harmony, Form, Reasoning, and Terminology. GPT-4o-mini scored the same responses using three prompting strategies: few-shot prompting with chain-of-thought reasoning (Fs+CoT), retrieval-augmented generation (RAG), and self-consistency based on five internal generations per administration (SC). Each strategy was administered three times with the model, prompt, rubric, and response held constant. Single-pass scores represented an operational scoring condition, whereas median aggregation across three runs was used to examine robustness. Agreement with teacher mean scores was evaluated using correlation, intraclass correlation, Krippendorff's alpha, quadratic weighted kappa, and scoring error indices. Fs+CoT showed the strongest agreement with teacher mean scores in both single-pass scoring and median aggregation. RAG showed systematic over-scoring, whereas SC produced highly repeatable scores but weaker individual-level agreement. Dimension-level analyses showed that scoring performance varied across rubric components, with Terminology generally showing weaker agreement than Reasoning. These findings indicate that GPT-4o-mini can generate stable scores for complex music analysis responses, but prompting strategies produce distinct scoring profiles. Operational use therefore requires strategy-specific calibration, dimension-level validation, and continued human oversight.
Problem

Research questions and friction points this paper is trying to address.

automated scoring
music analysis
open-ended responses
rubric-based assessment
AI validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

automated scoring
prompting strategies
GPT-4o-mini
music analysis
repeatability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Baicheng Lin
Department of Education, Sejong University, South Korea
L
Lingxi Jin
Department of Educational Technology, Ewha Womans University, South Korea
K
Kyung-Seok Min
Department of Education, Sejong University, South Korea