Attention Quantization for Tabular Foundation Models
研究针对表格基础模型的推理性能优化问题,提出一种将注意力机制中的查询、键和值量化为FP8的方法,并使用FP8矩阵乘法指令加速计算。
研究针对表格基础模型的推理性能优化问题,提出一种将注意力机制中的查询、键和值量化为FP8的方法,并使用FP8矩阵乘法指令加速计算。
研究探讨了表格基础模型通过单一真实表格自监督预训练实现强迁移的问题,提出任务中心和基于检索的视角来解释其泛化性能。
To address the low training efficiency, reliance on complex hyperparameter tuning, or large-scale architectures in lightweight time-series foundation models, this paper proposes SynthTS—a synthetic data generation and augmentation pipeline integrated with a causal input normalization mechanism. This combination enables, for the first time, rapid convergence and state-of-the-art (SOTA) performance for small models on dense forecasting tasks. Our 23M-parameter architecture trains efficiently on a single A100 GPU using next-token prediction loss and a hyperparameter-free training protocol. In medium- to long-term forecasting, it achieves uniformly lower MSE than existing methods—matching the performance of large industrial models—while remaining highly competitive in short-term forecasting. The core contribution lies in enabling efficient, low-cost, high-performance time-series foundation modeling under resource constraints, without neural architecture search or manual hyperparameter optimization.
To address the scalability limitations and accuracy bottlenecks in tabular data modeling, this paper introduces TabPFN-2.5, a next-generation foundation model. It is the first neural process architecture scaled to handle up to 50,000 samples and 2,000 features—achieving a 20× capacity increase over prior models. We propose a novel knowledge distillation engine that generates compact, high-performance surrogate models—including lightweight MLPs or tree ensembles—balancing accuracy and inference latency. On the TabArena benchmark, TabPFN-2.5 matches AutoGluon 1.4 for the first time and consistently outperforms XGBoost. It achieves a 100% win rate against default XGBoost on small-to-medium classification datasets, and maintains strong performance on large-scale data with 87% (classification) and 85% (regression) win rates—significantly surpassing tuned tree-based models.
研究针对表格基础模型的推理性能优化问题,提出一种将注意力机制中的查询、键和值量化为FP8的方法,并使用FP8矩阵乘法指令加速计算。
研究探讨了表格基础模型通过单一真实表格自监督预训练实现强迁移的问题,提出任务中心和基于检索的视角来解释其泛化性能。
To address the low training efficiency, reliance on complex hyperparameter tuning, or large-scale architectures in lightweight time-series foundation models, this paper proposes SynthTS—a synthetic data generation and augmentation pipeline integrated with a causal input normalization mechanism. This combination enables, for the first time, rapid convergence and state-of-the-art (SOTA) performance for small models on dense forecasting tasks. Our 23M-parameter architecture trains efficiently on a single A100 GPU using next-token prediction loss and a hyperparameter-free training protocol. In medium- to long-term forecasting, it achieves uniformly lower MSE than existing methods—matching the performance of large industrial models—while remaining highly competitive in short-term forecasting. The core contribution lies in enabling efficient, low-cost, high-performance time-series foundation modeling under resource constraints, without neural architecture search or manual hyperparameter optimization.
To address the scalability limitations and accuracy bottlenecks in tabular data modeling, this paper introduces TabPFN-2.5, a next-generation foundation model. It is the first neural process architecture scaled to handle up to 50,000 samples and 2,000 features—achieving a 20× capacity increase over prior models. We propose a novel knowledge distillation engine that generates compact, high-performance surrogate models—including lightweight MLPs or tree ensembles—balancing accuracy and inference latency. On the TabArena benchmark, TabPFN-2.5 matches AutoGluon 1.4 for the first time and consistently outperforms XGBoost. It achieves a 100% win rate against default XGBoost on small-to-medium classification datasets, and maintains strong performance on large-scale data with 87% (classification) and 85% (regression) win rates—significantly surpassing tuned tree-based models.