Why AI Detection Fails for Academic Integrity

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of commercial AI detectors in distinguishing between compliant, human-supervised AI-assisted editing and fully AI-generated text, a shortcoming that risks erroneous accusations of academic misconduct. By comparatively analyzing published English abstracts from four disciplines across two time periods (2013–2015 and 2023–2025) and employing leading detection tools such as Pangram and GPTZero alongside textual features—including long-token density and academic vocabulary density—the research quantifies, for the first time, a false-positive rate of 64%–80% for legitimate AI-edited content. Moreover, it demonstrates that AI-generated text subjected to human-like post-editing evades detection in over 96% of cases, reducing detection rates to below 4%. These findings critically challenge the validity of using AI detection scores as reliable evidence of academic dishonesty.
📝 Abstract
Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pangram/GPTZero). Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.
Problem

Research questions and friction points this paper is trying to address.

AI detection
academic integrity
AI editing
false positives
academic misconduct
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI detection failure
academic integrity
humanization evasion
LLM editing
proxy labeling
🔎 Similar Papers
2024-06-21Journal of Artificial Intelligence ResearchCitations: 6
J
Jonathan A. Karr Jr.
University of Notre Dame
G
Grigorii Khvatskii
University of Notre Dame
Ting Hua
Ting Hua
University of Notre Dame
Efficient learningCompressionReasoning
N
Nitesh V. Chawla
University of Notre Dame