Prediction-Oriented Subsampling from Data Streams
To address the challenge of balancing sampling efficiency and information preservation in offline learning under data stream settings, this paper proposes a prediction-oriented information-theoretic subsampling framework. Unlike conventional approaches that maximize input data entropy, our method guides sampling decisions by minimizing posterior uncertainty of downstream prediction tasks. It incorporates a lightweight model-aware mechanism to ensure sampling stability and computational tractability. Extensive experiments on time-series forecasting and anomaly detection demonstrate that the proposed method significantly outperforms existing information-theoretic baselines: it achieves an average 12.7% reduction in prediction error at equivalent sampling rates, while maintaining scalable computational overhead. The core contribution lies in the first explicit formulation of predictive uncertainty as a principled subsampling criterion—unifying theoretical interpretability with practical performance.