Constraint-Data-Value-Maximization: Utilizing Data Attribution for Effective Data Pruning in Low-Data Environments

📅 2026-05-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing Shapley value–based data pruning methods struggle to effectively retain critical samples in low-data regimes. This work proposes Constrained Data Value Maximization (CDVM), which introduces constrained optimization into data value–driven pruning for the first time. By formulating the problem as a constrained optimization task, CDVM simultaneously maximizes the overall influence of the retained dataset while limiting the excessive contribution of any individual test sample, thereby achieving a balance between global impact and local fidelity. Experiments on the OpenDataVal benchmark demonstrate that CDVM significantly outperforms existing approaches when retaining only a small fraction of data, offering both superior performance and computational efficiency.
📝 Abstract
Attributing model behavior to training data is an evolving research field. A common benchmark is data removal, which involves eliminating data instances with either low or high values, then assessing a model's performance trained on the modified dataset. Many existing studies leverage Shapley-based data values for this task. In this paper, we demonstrate that these data values are not optimally suited for pruning low-value data when only a limited amount of data remains. To address this limitation, we introduce the Constraint-Data-Value-Maximization (CDVM) approach, which effectively utilizes data attributions for pruning in low-data scenarios. By casting pruning as a constrained optimization that both maximizes total influence and penalizes excessive per-test contributions, CDVM delivers robust performance when only a small fraction of the data is retained. On the OpenDataVal benchmark, CDVM shows strong performance and competitive runtime.
Problem

Research questions and friction points this paper is trying to address.

data pruning
low-data environments
data attribution
Shapley values
data valuation
Innovation

Methods, ideas, or system contributions that make the work stand out.

data pruning
data attribution
constrained optimization
low-data regime
Shapley values
🔎 Similar Papers
2024-06-16International Conference on Learning RepresentationsCitations: 12
D
Danilo Brajovic
Fraunhofer IPA, Stuttgart, Germany; Institute of Industrial Manufacturing and Engineering IFF, University of Stuttgart, Germany
D
David A. Kreplin
Hochschule Heilbronn, Germany
M
Marco F. Huber
Fraunhofer IPA, Stuttgart, Germany; Institute of Industrial Manufacturing and Engineering IFF, University of Stuttgart, Germany