Beyond Participant-Level Cross-Validation: Reliable Inference for Longitudinal Machine Learning

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对纵向机器学习研究中由于数据分割不当导致的伪复制问题,提出了一种基于参与者标签置换并重运行整个工作流程的方法来校准结果。
📝 Abstract
Longitudinal sensing studies routinely collect thousands of windows from a few dozen participants. The records are numerous; the independent scientific units are not. When the outcome is defined per participant, this mismatch makes apparently precise findings vulnerable to pseudo-replication, to partition choice, and to the ordinary analytic flexibility of comparing several pipelines before reporting one. Splitting on participants prevents a person's records from straddling a split, but it does not calibrate the label-dependent workflow fold construction, preprocessing, tuning, calibration, and candidate selection that produced the reported number. We define a participant-level estimand and obtain an analysis-matched null by permuting participant labels and rerunning that entire workflow. In controlled simulation, window-level inference rejects in 70-80% of replicates when no effect exists and a window bootstrap rejects at the same rate; a participant bootstrap still rejects at 10-17%; the analysis-matched test holds 0.025-0.100 across cohorts of 20 to 80 participants. Freezing the selected pipeline instead of repeating the search inflates Type-I error to 0.240 with eight candidates, where repeating it holds 0.040. Applied to two public cohorts, wrist actigraphy (n=55) yields participant AUROC 0.928 with p=0.0050, a conclusion that persists under a scale-robust rank-pooled statistic and under a matched permutation null computed after excluding hospitalized participants (p=0.0089). Smartphone sensing (n=38, 7 positives) yields 0.636 and does not reject (p=0.1724) despite sufficient resolution, with sensitivity 0.143. The practical rule is narrow: every step that reads labels belongs inside the permuted analysis, and repeated records do not create additional independent participants.
Problem

Research questions and friction points this paper is trying to address.

Longitudinal Machine Learning
Pseudo-replication
Participant-level Cross-Validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

participant-level estimand
label permutation
analysis-matched null
Type-I error control
longitudinal machine learning