🤖 AI Summary
This study addresses the lack of a unified evaluation benchmark and insufficient attention to temporal interpretability and calibration in existing dropout prediction research within learning analytics. The authors construct the first multidimensional benchmark tailored for survival analysis, systematically comparing diverse models—including random survival forests, piecewise exponential additive models, parametric survival models, and neural survival models—under both dynamic weekly-granularity and continuous-time representations. Leveraging person-period data formatting and refit-free bootstrapping, they conduct a comprehensive assessment through a four-dimensional framework encompassing predictive performance, ablation, interpretability, and calibration. Results reveal that temporal behavioral features dominate predictive signals, whereas static background factors exert limited influence; random survival forests perform best under continuous-time settings, while piecewise exponential models show slight advantages in dynamic settings. Notably, models with high discriminative ability generally exhibit strong calibration.
📝 Abstract
Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heterogeneous protocols, prioritizing discrimination over temporal interpretability and calibration. This study introduces a survival-oriented benchmark for temporal dropout risk modelling using the Open University Learning Analytics Dataset (OULAD). Two harmonized arms are compared: a dynamic weekly arm, with models in person-period representation, and a comparable continuous-time arm, with an expanded roster of families -- tree-based survival, parametric, and neural models. The evaluation protocol integrates four analytical layers: predictive performance, ablation, explainability, and calibration. Results are reported within each arm separately, as a single cross-arm ranking is not methodologically warranted. Within the comparable arm, Random Survival Forest leads in discrimination and horizon-specific Brier scores; within the dynamic arm, Poisson Piecewise-Exponential leads narrowly on integrated Brier score within a tight five-family cluster. No-refit bootstrap sampling variability qualifies these positions as directional signals rather than absolute superiority. Ablation and explainability analyses converged, across all families, on a shared finding: the dominant predictive signal was not primarily demographic or structural, but temporal and behavioral. Calibration corroborated this pattern in the better-discriminating models, with the exception of XGBoost AFT, which exhibited systematic bias. These results support the value of a harmonized, multi-dimensional benchmark in Learning Analytics and situate dropout risk as a temporal-behavioral process rather than a function of static background attributes.