Can Large Language Models Differentiate Harmful from Argumentative Essays? Steps Toward Ethical Essay Scoring
This study addresses a critical gap in automated essay scoring (AES) systems and large language models (LLMs), which often fail to detect harmful content—such as racism or gender bias—in argumentative essays and may erroneously assign high scores to such texts. To tackle this issue, the work introduces an ethical dimension into AES evaluation for the first time, presenting the Harmful Essay Detection (HED) benchmark dataset specifically designed to assess the ability to identify harmful arguments. It also proposes a corresponding ethical evaluation framework. Experimental results demonstrate that prevailing LLMs and AES systems generally lack the capacity to distinguish between harmful and legitimate reasoning, revealing a significant blind spot in their moral judgment. This research establishes a foundational benchmark and offers a clear direction for developing ethically aware next-generation AES systems.