🤖 AI Summary
Current automated scoring and feedback for geometric construction tasks in online mathematics platforms lack personalization and formative support. Method: We propose a personalized automated scoring and feedback system leveraging GPT-4, integrating prompt engineering and few-shot learning. Structured prompts—built from teacher-annotated canonical student responses—enable dynamic generation of fine-grained, pedagogically aligned feedback and support iterative student revision. Contribution/Results: Deployed on the Algeomath platform, the system was evaluated with 79 middle-school students. Automated scores demonstrated high agreement with expert teacher judgments (Cohen’s κ = 0.86). Feedback significantly improved students’ error identification accuracy and task completion rates. These results validate the system’s effectiveness and educational viability for scalable, formative assessment in geometry learning.
📝 Abstract
As personalized learning gains increasing attention in mathematics education, there is a growing demand for intelligent systems that can assess complex student responses and provide individualized feedback in real time. In this study, we present a personalized auto-grading and feedback system for constructive geometry tasks, developed using large language models (LLMs) and deployed on the Algeomath platform, a Korean online tool designed for interactive geometric constructions. The proposed system evaluates student-submitted geometric constructions by analyzing their procedural accuracy and conceptual understanding. It employs a prompt-based grading mechanism using GPT-4, where student answers and model solutions are compared through a few-shot learning approach. Feedback is generated based on teacher-authored examples built from anticipated student responses, and it dynamically adapts to the student's problem-solving history, allowing up to four iterative attempts per question. The system was piloted with 79 middle-school students, where LLM-generated grades and feedback were benchmarked against teacher judgments. Grading closely aligned with teachers, and feedback helped many students revise errors and complete multi-step geometry tasks. While short-term corrections were frequent, longer-term transfer effects were less clear. Overall, the study highlights the potential of LLMs to support scalable, teacher-aligned formative assessment in mathematics, while pointing to improvements needed in terminology handling and feedback design.