Towards Scalable Measurement of Durable Skills

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用大语言模型与人类互动,解决耐用技能难以测量的问题,实现了在保持生态效度的同时保证心理测量的严谨性。
📝 Abstract
Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an "Executive LLM" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.
Problem

Research questions and friction points this paper is trying to address.

durable skills
measurement
educational curricula
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Durable Skills Assessment
Ecological Validity
Psychometric Rigor
Executive LLM
💼 Related Jobs
No related jobs found.
Amir Globerson
Amir Globerson
Tel Aviv University
Machine LearningGraphical ModelsOptimizationNatural Language Processing
A
Amy Keeling
Google Research
A
Anisha Choudhury
Google Research
Anna Iurchenko
Anna Iurchenko
Google
UX DesignHuman Computer InteractionDesign ThinkingBehavioral ChangeHealth
A
Aviad Segal
OpenMic
A
Avinatan Hassidim
Google Research
A
Ayça Çakmakli
Google Research
B
Ben Gomes
Google Research
B
Benn Witt
Google Research
C
Cathy Cheung
Google Research
C
Cristine Legare
The University of Texas at Austin
D
Diana Akrong
Google Research
E
Eliad Carmi
OpenMic
E
Elisabeth Bauer
Google Research
Gal Elidan
Gal Elidan
Hebrew University / Google
Machine LearningGraphical ModelsStructure LearningCopulas
H
Hadas Gelbart
OpenMic
H
Hairong Mu
Google Research
Katherine Chou
Katherine Chou
Google
MLHealthGraphics
L
Lev Borovoi
Google Research
N
Nir Kerem
Google Research
N
Niv Efron
Google Research
N
Noa Kerrem Gilo
Google Research
P
Preeti Singh
Google Research
R
Rajvi Kapadia
Google Research
R
Rena Levitt
Google Research