Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses ongoing validity concerns regarding the direct transfer of human assessments to AI evaluation by examining human-AI construct equivalence from a psychometric perspective. Integrating exploratory factor analysis, consistency testing, and resampling methods, we systematically compared latent structures between humans and large language models in educational assessments. The results provide the first empirical evidence of significantly divergent factor structures in chemistry and quantitative reasoning tasks, demonstrating construct nonequivalence. These findings challenge the prevailing paradigm of interpreting AI capabilities through human normative standards. Instead, this work posits that assessment transfer must be predicated on latent structural similarity, thereby offering a critical theoretical foundation for developing scientifically rigorous AI capability evaluation frameworks.
📝 Abstract
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.
Problem

Research questions and friction points this paper is trying to address.

Assessment Validity
Latent Structure
Large Language Models
Construct Equivalence
Psychometrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Structure Analysis
Assessment Validity
Factor Congruence
LLM Evaluation
Construct Equivalence
💼 Related Jobs
No related jobs found.
A
Alona Strugatski
Weizmann Institute of Science
L
Licol Zeinfeld
Weizmann Institute of Science
Giora Alexandron
Giora Alexandron
Associate Professor, Weizmann Institute of Science
AI in EducationLearning AnalyticsEducational Data MiningAI Education