MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
Industrial-scale tabular data often confronts challenges such as high dimensionality, extensive missing values, and scarce annotations, with a notable absence of a general-purpose self-supervised pretraining framework. This work proposes MaskTab, a unified pretraining approach that employs learnable missing tokens to distinguish between structural and random missingness, integrates a dual-path hybrid supervision architecture to jointly optimize masked reconstruction and downstream task objectives, and incorporates a Mixture-of-Experts (MoE)-enhanced loss with knowledge distillation. Evaluated on industrial benchmarks, MaskTab achieves substantial performance gains (AUC +5.04%, KS +8.28%). Moreover, the distilled lightweight model retains strong performance under stringent latency and interpretability constraints (AUC +2.55%, KS +4.85%) while demonstrating enhanced robustness to distribution shifts.