Calibrating Generative AI to Produce Realistic Essays for Data Augmentation
This study addresses the performance bottleneck in machine learning–based automated essay scoring systems caused by limited training data by proposing the use of large language models to generate synthetic student essays for data augmentation. The work presents the first empirical evaluation of three prompting strategies—“next-sentence prediction,” “sentence-level prompting,” and “25-shot exemplars”—systematically comparing their effectiveness in generating text that preserves original essay quality and exhibits human-like authenticity. Results indicate that the next-sentence prediction strategy achieves the highest scoring consistency and, alongside sentence-level prompting, best retains the quality of the source essays. Moreover, texts generated via next-sentence prediction and the 25-shot approach demonstrate the greatest authenticity. This research provides both effective strategies and empirical evidence supporting the use of synthetic data augmentation in automated essay scoring.