Trading off Consistency and Dimensionality of Convex Surrogates for the Mode
To address the intractability of surrogate loss optimization in large-scale multiclass classification caused by high-dimensional embeddings, this paper proposes a low-dimensional convex polyhedral embedding framework, establishing theoretical trade-offs among embedding dimension, consistency regions, and data distribution assumptions. First, it rigorously proves that “hallucination”—i.e., spurious class predictions—necessarily occurs when the embedding dimension is less than $n-1$. Second, under a low-noise assumption, it derives a verifiable consistency criterion. Third, it designs structured embeddings—including hypercubes and permutahedra—that achieve dimensionality reductions from $2^d$ to $d$ and from $d!$ to $d$, respectively. Finally, in the multiple-instance learning setting, it shows that full simplex consistency is guaranteed with only $n/2$ embedding dimensions and proves the existence of consistent subsets around any point-mass distribution.