Beyond "AI Language": The case for the idiolectal nature of LLM output
This study challenges the conventional view that large language model (LLM) outputs constitute a homogeneous “AI language,” arguing instead that individual LLMs exhibit stable, idiolectal linguistic styles akin to human speakers. Drawing on text corpora generated by two distinct LLMs from 2024 and 2026, the research employs computational linguistic descriptors and stylometric principal component analysis (PCA) to empirically assess stylistic variation. Findings reveal significant and consistent differences between models—even under identical prompts—with notable disparities such as a 25-fold variation in contraction usage frequency. These results robustly support the existence of model-specific idiolects, offering novel insights for research in linguistic variation, authorship attribution, and forensic text analysis.