Leveraging Language Models for Analyzing Longitudinal Experiential Data in Education

📅 2025-03-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses early academic trajectory prediction for STEM students, tackling challenges inherent in educational longitudinal data—including high missingness rates, limited sample sizes, and multimodal temporal variability. We propose a comprehensive data augmentation framework tailored to educational experience data, integrating statistical imputation, few-shot augmentation, and task-instruction embedding to support both encoder-decoder and decoder-only pretrained language models. To enhance modeling of heterogeneous temporal patterns, we introduce context-aware embeddings and instruction tuning. Experimental results demonstrate substantial improvements over baselines in multimodal fusion performance and robustness to missing data. However, the findings also reveal that current language models rely more heavily on high-level statistical regularities than on explicit temporal logic—a critical empirical insight into the interpretability limitations and applicability boundaries of AI in education.

Technology Category

Application Category

📝 Abstract
We propose a novel approach to leveraging pre-trained language models (LMs) for early forecasting of academic trajectories in STEM students using high-dimensional longitudinal experiential data. This data, which captures students' study-related activities, behaviors, and psychological states, offers valuable insights for forecasting-based interventions. Key challenges in handling such data include high rates of missing values, limited dataset size due to costly data collection, and complex temporal variability across modalities. Our approach addresses these issues through a comprehensive data enrichment process, integrating strategies for managing missing values, augmenting data, and embedding task-specific instructions and contextual cues to enhance the models' capacity for learning temporal patterns. Through extensive experiments on a curated student learning dataset, we evaluate both encoder-decoder and decoder-only LMs. While our findings show that LMs effectively integrate data across modalities and exhibit resilience to missing data, they primarily rely on high-level statistical patterns rather than demonstrating a deeper understanding of temporal dynamics. Furthermore, their ability to interpret explicit temporal information remains limited. This work advances educational data science by highlighting both the potential and limitations of LMs in modeling student trajectories for early intervention based on longitudinal experiential data.
Problem

Research questions and friction points this paper is trying to address.

Forecasting academic trajectories using longitudinal experiential data
Handling missing values and limited dataset size in educational data
Enhancing language models' understanding of temporal dynamics in student behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pre-trained LMs forecast academic trajectories
Data enrichment handles missing values
LMs integrate multimodal data resiliently
A
Ahatsham Hayat
Electrical and Computer Engineering, University of Nebraska-Lincoln
Bilal Khan
Bilal Khan
Professor of Population Health & Computer Science
Population HealthComputer ScienceMathematicsSociology
M
Mohammad Rashedul Hasan
Electrical and Computer Engineering, University of Nebraska-Lincoln