The Art of Misclassification: Too Many Classes, Not Enough Points
This paper addresses the inherent difficulty of classification under high-class-count, low-sample-size regimes. We propose an information-theoretic “class separability” metric grounded in entropy, which formally characterizes the irreducible inter-class overlap and uncertainty intrinsic to a dataset in feature space. Unlike prior measures, our metric is model-agnostic and sample-size-independent, enabling derivation of a fundamental theoretical upper bound on classification accuracy—i.e., a performance ceiling that no classifier can surpass. Leveraging entropy analysis and uncertainty modeling, we establish a tight generalization bound and empirically validate that this bound aligns closely with human perception of ambiguous decision boundaries. Our core contribution is the formal definition and quantification of the intrinsic solvability of a classification task, thereby providing a principled theoretical benchmark for algorithm design, model selection, and dataset evaluation.