Provenance Networks: End-to-End Exemplar-Based Explainability

📅 2025-10-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Deep learning models suffer from poor interpretability, frequent hallucinations, and difficulty tracing predictions back to training samples—undermining their robustness and trustworthiness. To address this, we propose TraceNet, an architecture that intrinsically embeds example-based interpretability via end-to-end prediction-to-sample association. Our core innovation is a learnable weighted k-nearest neighbors (k-NN) mechanism that dynamically retrieves and weights supportive training instances in the feature space, jointly optimizing both the primary task loss and an explicit interpretability objective. TraceNet enables training-data provenance, label-noise detection, improved robustness to input perturbations, and explicit generation grounding. Experiments on medium-scale benchmarks demonstrate substantial gains in decision transparency and yield novel insights into the interplay between memorization and generalization.

Technology Category

Application Category

📝 Abstract
We introduce provenance networks, a novel class of neural models designed to provide end-to-end, training-data-driven explainability. Unlike conventional post-hoc methods, provenance networks learn to link each prediction directly to its supporting training examples as part of the model's normal operation, embedding interpretability into the architecture itself. Conceptually, the model operates similarly to a learned KNN, where each output is justified by concrete exemplars weighted by relevance in the feature space. This approach facilitates systematic investigations of the trade-off between memorization and generalization, enables verification of whether a given input was included in the training set, aids in the detection of mislabeled or anomalous data points, enhances resilience to input perturbations, and supports the identification of similar inputs contributing to the generation of a new data point. By jointly optimizing the primary task and the explainability objective, provenance networks offer insights into model behavior that traditional deep networks cannot provide. While the model introduces additional computational cost and currently scales to moderately sized datasets, it provides a complementary approach to existing explainability techniques. In particular, it addresses critical challenges in modern deep learning, including model opaqueness, hallucination, and the assignment of credit to data contributors, thereby improving transparency, robustness, and trustworthiness in neural models.
Problem

Research questions and friction points this paper is trying to address.

Providing end-to-end training-data-driven explainability for neural models
Linking predictions directly to supporting training examples for interpretability
Addressing model opaqueness, hallucination, and data contributor credit assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Provenance networks link predictions to training examples
Model embeds interpretability directly into architecture design
Jointly optimizes primary task and explainability objectives
🔎 Similar Papers
No similar papers found.
A
Ali Kayyam
BrainChip Inc., 23041 Avenida De La Carlota, Suite 250, Laguna Hills, CA 92653, USA
A
Anusha Madan Gopal
BrainChip Inc., 23041 Avenida De La Carlota, Suite 250, Laguna Hills, CA 92653, USA
M
M. Anthony Lewis
BrainChip Inc., 23041 Avenida De La Carlota, Suite 250, Laguna Hills, CA 92653, USA