SteLLA: A Structured Grading System Using LLMs with RAG

📅 2024-12-15
🏛️ BigData Congress [Services Society]
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses fairness deficiencies and poor interpretability of feedback in Large Language Models (LLMs) for Automated Short Answer Grading (ASAG). We propose a novel grading framework integrating Retrieval-Augmented Generation (RAG) with structured question-answering assessment. Our method explicitly models instructor-provided reference answers and rubrics, aligning domain-specific educational knowledge via structured prompt engineering to enable knowledge-component–level fine-grained scoring and natural-language feedback. Crucially, we are the first to deeply embed RAG within a structured assessment pipeline, substantially improving LLMs’ fidelity to grading criteria and subject-matter semantics. Evaluated on authentic university-level biology exam data, our system achieves high inter-rater agreement with human graders (Cohen’s κ > 0.8) and supports comprehensive, tiered scoring across all knowledge components. Error analysis reveals strong factual recognition but notable over-inference tendencies—establishing a new paradigm for trustworthy AI-assisted educational assessment.

Technology Category

Application Category

📝 Abstract
Large Language Models (LLMs) have shown strong general capabilities in many applications. However, how to make them reliable tools for some specific tasks such as automated short answer grading (ASAG) remains a challenge. We present SteLLA (Structured Grading System Using LLMs with RAG) in which a) Retrieval Augmented Generation (RAG) approach is used to empower LLMs specifically on the ASAG task by extracting structured information from the highly relevant and reliable external knowledge based on the instructor-provided reference answer and rubric, b) an LLM performs a structured and question-answering-based evaluation of student answers to provide analytical grades and feedback. A real-world dataset which contains students’ answers in an exam was collected from a college-level Biology course. Experiments show that our proposed system can achieve substantial agreement with the human grader while providing break-down grades and feedback on all the knowledge points examined in the problem. A qualitative and error analysis of the feedback generated by GPT4 shows that GPT4 is good at capturing facts while may prone to inferring too much implication from the given text in the grading task which provides insights into the usage of LLMs in the ASAG system.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Educational Assessment
Fairness and Accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

RAG Technology
Large Language Models
Automated Essay Grading
💼 Related Jobs
No related jobs found.
Hefei Qiu
Hefei Qiu
Assistant Professor @ Fitchburg State University
B
Brian White
Department of Computer Science, University of Massachusetts Boston, 100 Morrissey Blvd, Boston, MA 02125
A
Ashley Ding
Chantilly High School, 4201 Stringfellow Rd, Chantilly, VA 20151
R
Reinaldo Costa
Department of Computer Science, University of Massachusetts Boston, 100 Morrissey Blvd, Boston, MA 02125
A
Ali Hachem
Department of Computer Science, University of Massachusetts Boston, 100 Morrissey Blvd, Boston, MA 02125
W
Wei Ding
Department of Computer Science, University of Massachusetts Boston, 100 Morrissey Blvd, Boston, MA 02125
P
Ping Chen
Department of Computer Science, University of Massachusetts Boston, 100 Morrissey Blvd, Boston, MA 02125