Paraphrase Types Elicit Prompt Engineering Capabilities
This study investigates how linguistic dimensions of prompt formulation—morphology, syntax, and lexis—affect large language model (LLM) task performance. Method: We conduct controlled experiments across 120 diverse tasks using five LLMs, rigorously matching confounding factors such as prompt length and lexical diversity. This enables the first fine-grained, linguistically grounded attribution analysis of prompt rewriting. Contribution/Results: Semantic-preserving rewrites at morphological and lexical levels yield the largest performance gains, revealing LLMs’ high sensitivity to surface-form variation. We propose a multidimensional prompt rewriting generation framework that integrates cross-model behavioral comparison and median gain statistics. Evaluated on Mixtral 8x7B and LLaMA-3-8B, it achieves median task performance improvements of +6.7% and +5.5%, respectively—demonstrating that linguistically informed rewriting systematically enhances prompt robustness and effectiveness.