GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决表格基础模型的高延迟和内存成本问题,提出GEAR框架,通过两阶段蒸馏方法将其转化为轻量级预测器,以减少部署成本并提高性能。
📝 Abstract
Tabular foundation models (TFMs) achieve strong performance through in-context learning, but context-dependent inference imposes substantial latency and memory costs, hindering large-scale deployment. We propose GEAR (\emph{Generative Expansion and Real Anchoring}), a modular two-stage framework that distills TFMs into lightweight MLP or tree-based predictors that can be deployed on commodity CPUs. Stage 1 uses synthetic covariates solely as teacher-query locations and trains the student on soft TFM targets, expanding coverage beyond observed rows. Stage 2 re-anchors the student to the target distribution using real labels and out-of-fold teacher predictions, whitch avoids self-labeling leakage. We further derive a risk certificate characterizing the trade-off between generated-query volume and generator fidelity. Experiments on TALENT and TabArena demonstrate the broad applicability of GEAR. Two-stage MLPs outperform supervised MLPs by 1.81--2.00 AUC points on binary tasks and 1.19--1.35 points on multiclass tasks, with additional gains over real-data-only distillation of 1.76--2.19 and 2.09--2.40 points, respectively. On binary tasks, the gains also transfer to LightGBM and XGBoost, and all three student families outperform CatBoost, the strongest non-TFM baseline, in mean AUC. Ablations show gains beyond longer training or alternative warm starts, greater stability from staged than mixed optimization, and generator-dependent diminishing returns as query volume increases. Finally, GEAR reduces median inference time by 57--2866 times and peak prediction memory by 1.9--3.3 times, while retaining higher AUC than matched supervised baselines.
Problem

Research questions and friction points this paper is trying to address.

Tabular Foundation Models
In-context Learning
Latency
Memory Costs
Large-scale Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Expansion and Real Anchoring
two-stage distillation
tabular foundation models
synthetic covariates
risk certificate
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qi Qin
Renmin University of China, Center for Applied Statistics and School of Statistics, Beijing, China
Jiajie Zhu
Jiajie Zhu
Macquarie University
recommender systemscross-domain recommendation
D
Dali Chen
Nanjing University, Nanjing, China
Yuzhao Zhang
Yuzhao Zhang
Ant Digital Technologies, Ant Group, Hangzhou, China
J
Jia-Xing Han
Ant Digital Technologies, Ant Group, Hangzhou, China
Y
Yu Su
Ant Digital Technologies, Ant Group, Hangzhou, China
P
Peng Zhang
Ant Digital Technologies, Ant Group, Hangzhou, China
Ying Yan
Ying Yan
Microsoft Research
Big Data Management
Y
Yifan Sun
Renmin University of China, Center for Applied Statistics and School of Statistics, Beijing, China