Distilling Foundation Models for Agentic What-If Reasoning:Cost, Latency, and Governance in a Hybrid LLM+SLM Architecture

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在混合LLM+SLM架构中蒸馏表格基础模型,解决了高推理延迟问题,大幅压缩参数同时保持了高性能。
📝 Abstract
Tabular foundation models deliver strong zero-training predictive performance via in-context learning, but their high inference latency makes them impractical as hot-path decision backends in interactive agentic loops. We distill a TabPFN teacher into a compact feed-forward student across a business-decision simulation on UCI Adult and five OpenML benchmarks: the classification head compresses 53.2M parameters to 8,546 (6,220x); the deployed two-head loan pipeline compresses 111.4M parameters to 17,059 (6,532x). The student retains 95.4-100.5% accuracy and 96.8-100.0% AUC, with the lowest accuracy retention on credit-g at 95.4%; an alpha = 0 hard-label control shows that the teacher's soft targets provide a 2.1-7.0 AUC point gain.
Problem

Research questions and friction points this paper is trying to address.

Tabular Foundation Models
Inference Latency
Interactive Agentic Loops
Innovation

Methods, ideas, or system contributions that make the work stand out.

model distillation
low-latency inference
compact model
in-context learning
high accuracy retention
S
Sourish Dey
Machine Learning, SumUp, Berlin
A
Aditya Kumar
Institute of Physics, Johannes Gutenberg-Universität Mainz, 55128 Mainz, Germany