Improving Model Classification by Optimizing the Training Dataset
This work addresses the limitation of conventional dataset optimization methods, which overly rely on loss approximation and neglect downstream classification metrics. We propose a systematic coreset construction framework explicitly designed to enhance classification performance. Methodologically, it integrates deterministic sampling, class-level importance weighting, and active learning–based refinement, optimizing directly for task-oriented metrics such as the F1 score—thereby overcoming the narrow focus of sensitivity-based sampling on gradient or loss approximation alone. Our key innovation lies in the joint design of importance sampling and tunable parameters to enable end-to-end optimization of training data quality. Extensive experiments across multiple benchmark datasets and classifiers demonstrate that our approach significantly outperforms standard coreset baselines and even full-data training, achieving superior classification accuracy while maintaining high training efficiency.