Institution profile

Dataheroes

Industry researchnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Improving Model Classification by Optimizing the Training Dataset

Jul 22, 2025

This work addresses the limitation of conventional dataset optimization methods, which overly rely on loss approximation and neglect downstream classification metrics. We propose a systematic coreset construction framework explicitly designed to enhance classification performance. Methodologically, it integrates deterministic sampling, class-level importance weighting, and active learning–based refinement, optimizing directly for task-oriented metrics such as the F1 score—thereby overcoming the narrow focus of sensitivity-based sampling on gradient or loss approximation alone. Our key innovation lies in the joint design of importance sampling and tunable parameters to enable end-to-end optimization of training data quality. Extensive experiments across multiple benchmark datasets and classifiers demonstrate that our approach significantly outperforms standard coreset baselines and even full-data training, achieving superior classification accuracy while maintaining high training efficiency.

0 citationsRead paper
Recent publications

Latest Papers

Improving Model Classification by Optimizing the Training Dataset

Jul 22, 2025

This work addresses the limitation of conventional dataset optimization methods, which overly rely on loss approximation and neglect downstream classification metrics. We propose a systematic coreset construction framework explicitly designed to enhance classification performance. Methodologically, it integrates deterministic sampling, class-level importance weighting, and active learning–based refinement, optimizing directly for task-oriented metrics such as the F1 score—thereby overcoming the narrow focus of sensitivity-based sampling on gradient or loss approximation alone. Our key innovation lies in the joint design of importance sampling and tunable parameters to enable end-to-end optimization of training data quality. Extensive experiments across multiple benchmark datasets and classifiers demonstrate that our approach significantly outperforms standard coreset baselines and even full-data training, achieving superior classification accuracy while maintaining high training efficiency.

0 citationsRead paper