Black-Box Knowledge Transfer across Distinct Feature Sets

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of transferring knowledge from pretrained black-box models when input features reside in a space inconsistent with the model’s expected domain. The authors propose a two-stage neural network approach that decomposes the target regression function into transferable and non-transferable components. The transferable part is estimated via unsupervised alignment using unlabeled cross-space feature pairs, while the non-transferable component is learned from a small set of labeled data. This method achieves, for the first time, knowledge transfer from one or multiple black-box models to heterogeneous feature spaces. Theoretically, it yields a risk upper bound strictly tighter than that of minimax estimators relying solely on labeled data, and adapts automatically to diverse scenarios. Empirical results demonstrate significant improvements in prediction accuracy on both synthetic and real-world datasets.
πŸ“ Abstract
Pre-trained black-box predictive functions encode knowledge distilled from massive datasets and extensive computation. However, when the available input features differ from those the black box expects, direct use is infeasible. We introduce a method for transferring predictive knowledge from the black box to a new, heterogeneous input space. Our approach decomposes the target regression function into a transferable component, which the black box can inform, and a non-transferable component, which captures information unique to the new space. We propose a two-step neural network procedure, estimating the transferable component from abundant unlabeled feature pairs that bridge the two input spaces and the non-transferable component from limited labels. We derive prediction risk bounds that improve on those of a non-transfer alternative when the non-transferable component is small or smooth, and the procedure adapts to either case. Under additional conditions, the worst-case risk of our estimator is of strictly smaller polynomial order than the minimax risk of estimation from the labeled data alone. We extend the framework to multiple black boxes, each on its own input space, and show that aggregation can reduce prediction error relative to the best single black box. Simulated and real data demonstrate the practical value of the method.
Problem

Research questions and friction points this paper is trying to address.

Black-box knowledge transfer
heterogeneous input spaces
feature mismatch
predictive knowledge transfer
transfer learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

black-box knowledge transfer
heterogeneous feature spaces
two-step neural estimation
prediction risk bounds
model aggregation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
O
Oh-Ran Kwon
Department of Statistics, The Ohio State University
D
Daeyoung Ham
Department of Statistics and Data Science, University of Texas at San Antonio