Score
Designing, validating, and analyzing measurement instruments (scales, assessments) to quantify latent constructs—ensuring reliability, validity, and comparability across experimental conditions and populations for constructs like embodiment, privacy perceptions, or acceptance.
Existing research frequently suffers from model misspecification of formative constructs, and the absence of a consensus-based validation methodology leads scholars to erroneously apply reflective measurement frameworks, thereby compromising construct validity. Method: This paper introduces the first dedicated, multi-stage validation framework for formative constructs, integrating systematic literature review, descriptive statistics, multicollinearity diagnostics, and formative-model-specific tests to rigorously distinguish formative (causal) from reflective (effect) measurement logic. Contribution/Results: The framework ensures both theoretical rigor and practical feasibility, substantially enhancing the psychometric soundness and statistical integrity of formative indicators. It provides a reproducible, defensible methodological pathway for scale development and construct validation, directly addressing longstanding measurement challenges in behavioral and social science research.
Traditional structural equation modeling (SEM) relies predominantly on reflective latent variable specifications, limiting its capacity to flexibly represent composite constructs—linear combinations of observed indicators. Existing compositional modeling approaches either compromise core SEM functionalities (e.g., overall model fit assessment, missing data handling, multi-group comparison) or inflate model complexity via auxiliary latent variables. Method: We propose the first SEM framework that unifies composites and latent variables within a single covariance structure model, leveraging maximum likelihood and generalized least squares estimation to directly specify the implied covariance matrix incorporating composites. Contribution/Results: Our approach eliminates the need for redundant latent variables while fully preserving SEM’s diagnostic and inferential capabilities—including fit evaluation, standard error estimation, and hypothesis testing. It significantly enhances expressive power and analytical flexibility for hybrid constructs (reflective + formative) and extends SEM’s applicability to more complex theoretical models.
HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.
Large language models (LLMs) are increasingly deployed in psychological research—as tools, targets of assessment, and cognitive models—yet recent evidence reveals severe measurement unreliability: factor structures of personality traits collapse, moral judgments reverse with minor punctuation changes, and theory-of-mind performance fluctuates dramatically under syntactic rephrasing. These “measurement ghosts” reflect statistical artifacts rather than substantive phenomena, threatening construct validity. Method: We propose the first validity-driven, six-stage workflow integrating psychometric principles and causal inference frameworks, dynamically calibrating validation rigor to research objectives and systematically governing the entire LLM psychology research lifecycle. Our approach includes construct validity verification, computational confound control, modeling of non-independent observations, and transparent experimental design. Contribution/Results: Applied to assessing “LLM selfhood,” our framework successfully disentangles genuine computational phenomena from measurement artifacts, establishing a reproducible empirical paradigm for AI psychology.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
This study addresses the frequent neglect of local cultural perspectives in existing automated evaluations of AI-generated images, particularly regarding “cultural appropriateness.” It introduces a novel evaluation framework that deeply integrates diverse community participation from the outset, collaborating with blind and visually impaired individuals in the UK and residents of Kerala and Tamil Nadu in India to systematically translate lived cultural experiences and community concerns into actionable assessment dimensions. Leveraging multimodal large language models as judges (LLM-as-a-judge), the approach operationalizes community consensus into structured scoring rules, enabling automated evaluation of cultural appropriateness. The work not only establishes a conceptual framework grounded in community values and demonstrates its feasibility but also exposes critical limitations in current AI models’ understanding of cultural context.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
This study addresses a critical yet previously unrecognized issue in observational causal inference: measurement-induced confounding, wherein latent variables—such as motivation or self-efficacy—are imperfectly measured, leading to biased estimates of adjusted causal effects. The authors formally identify and name this problem, moving beyond conventional two-stage adjustment approaches. They propose a novel Bayesian joint estimation framework that simultaneously models the latent variable’s measurement structure, the treatment assignment mechanism, and the potential outcomes model. This integrated approach effectively corrects bias in average treatment effect estimation and restores the nominal coverage of uncertainty intervals, thereby substantially enhancing the reliability of causal inferences drawn from observational data with error-prone proxies for unobserved confounders.
This study addresses the distortion in reliability reporting common in psychological research due to selective computation. For the first time, it applies a unified marginal reliability estimator across 889 psychometric datasets from the Item Response Warehouse, enabling standardized reliability assessment of cognitive tests, clinical screening instruments, and personality and attitude scales. Findings reveal that, depending on the reliability definition used, between 30% and 52% of datasets fall below the conventional .80 reliability threshold, with observed variability primarily attributable to genuine differences rather than estimation noise. The work underscores the substantial impact of reliability definition on substantive conclusions and advocates for more transparent and consistent reporting standards in psychometric practice.
This study addresses the lack of effective instruments for assessing computational empowerment and self-beliefs among adolescents engaged in constructing and deconstructing AI/ML systems. It develops and validates a six-factor self-belief measurement model tailored to youth, encompassing creative expression, problem-solving self-belief, auditing self-efficacy, interest in auditing, beliefs about design justice, and perceived value of learning AI/ML. Confirmatory factor analysis of survey data from 124 adolescents supports the validity of the proposed six-factor structure and reveals significant positive associations between design justice beliefs and problem-solving self-belief, auditing self-efficacy, and creative expression. By integrating dimensions of construction, deconstruction, and design justice, this work offers a novel instrument for evaluating adolescent AI literacy.