Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation

📅 2026-05-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current holistic evaluation approaches struggle to distinguish between structural replication and fidelity to user intent in large language model outputs. This work proposes a dimension-level intent fidelity assessment framework that, through structured prompt ablation, human evaluation, and weight perturbation, reveals for the first time a systematic divergence between structural fidelity and intent fidelity. Experiments show that among high-scoring outputs in both Chinese and English, 25.7% and 58.6%, respectively, exhibit deficiencies in dimensional intent alignment. Moreover, dimension-level scores demonstrate significantly higher agreement with human judgments than holistic scores, offering a more precise reflection of output quality flaws.
📝 Abstract
Holistic evaluation scores capture overall output quality but do not distinguish whether a model reproduced the structural form of a user's request from whether it preserved the user's specific intent. We propose a dimension-level intent fidelity evaluation framework, applied here through a structured prompt ablation study across 2,880 outputs spanning three languages, three task domains, and six LLMs, that separately measures structural recovery and intent fidelity for each semantic dimension. This framework reveals a systematic structural-fidelity split: among Chinese-language outputs with complete paired scores, 25.7% received perfect holistic alignment scores (GA=5) while exhibiting measurable dimensional intent deficits; among English-language outputs, this proportion rose to 58.6%. Human evaluation confirmed that these split-zone outputs represent genuine quality deficits and that dimensional fidelity scores track human judgements more reliably than holistic scores do. A public-private decomposition of 2,520 ablation cells characterises when models successfully compensate for missing intent and when they fail, while proxy annotation distinguishes prior inferability from default recoverability. A weight-perturbation experiment shows that moderate misalignment is typically absorbed, whereas severe dimensional inversion is consistently harmful. These findings demonstrate that dimension-level intent fidelity evaluation is a necessary complement to holistic assessment when evaluating LLM outputs for user-specific tasks.
Problem

Research questions and friction points this paper is trying to address.

intent fidelity
structured prompt ablation
dimension-level evaluation
large language models
holistic evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

intent fidelity
structured prompt ablation
dimension-level evaluation
structural-fidelity split
LLM output assessment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gang Peng
Huizhou Lateni AI Technology Co., Ltd., Huizhou, China; Huizhou University, Huizhou, China