When Imbalance Comes Twice: Active Learning under Simulated Class Imbalance and Label Shift in Binary Semantic Segmentation

πŸ“… 2026-01-08
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the degradation of active learning performance in binary semantic segmentation caused by the coexistence of class imbalance and label shift. For the first time, it systematically simulates both challenges jointly on open-source datasets to evaluate the effectiveness of three active learning strategies: random sampling, entropy maximization, and core-set selection. Experimental results demonstrate that entropy-based and core-set methods remain robust under severe class imbalance; however, strong label shift significantly impairs their performance. By revealing distinct behavioral patterns of these strategies under compound distribution shifts, this work provides critical insights for deploying active learning in real-world scenarios where multiple data biases may co-occur.

Technology Category

Application Category

πŸ“ Abstract
The aim of Active Learning is to select the most informative samples from an unlabelled set of data. This is useful in cases where the amount of data is large and labelling is expensive, such as in machine vision or medical imaging. Two particularities of machine vision are first, that most of the images produced are free of defects, and second, that the amount of images produced is so big that we cannot store all acquired images. This results, on the one hand, in a strong class imbalance in defect distribution and, on the other hand, in a potential label shift caused by limited storage. To understand how these two forms of imbalance affect active learning algorithms, we propose a simulation study based on two open-source datasets. We artificially create datasets for which we control the levels of class imbalance and label shift. Three standard active learning selection strategies are compared: random sampling, entropy-based selection, and core-set selection. We demonstrate that active learning strategies, and in particular the entropy-based and core-set selections, remain interesting and efficient even for highly imbalanced datasets. We also illustrate and measure the loss of efficiency that occurs in the situation a strong label shift.
Problem

Research questions and friction points this paper is trying to address.

active learning
class imbalance
label shift
semantic segmentation
binary classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

active learning
class imbalance
label shift
semantic segmentation
simulation study
πŸ’Ό Related Jobs
No related jobs found.
J
Julien Combes
Michelin, Clermont-Ferrand, 63000, France
A
Alexandre Derville
Michelin, Clermont-Ferrand, 63000, France
J
Jean-FranΓ§ois Coeurjolly
Laboratoire Jean Kuntzmann, Grenoble, 38000, France