Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
This study addresses the threat posed by high-quality AI-generated essays to the authenticity of writing assessment and the limited generalization capability of existing detectors in cross–large language model (LLM) scenarios. It presents the first systematic evaluation of mainstream AI text detectors on essays generated by multiple LLMs, constructing a multi-source dataset based on publicly available GRE prompts to conduct an empirical analysis of cross-model generalization. The findings reveal a significant performance drop in current detectors when applied across different LLMs. Building on these insights, the work proposes actionable retraining strategies and responsible deployment guidelines to enhance the robustness and practical utility of detection tools in real-world educational settings.