Sparse Prototype Code Underlies Classification and Prediction Across Modalities

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the interpretability challenges of high-dimensional neural representations and unclear cross-modal classification mechanisms by proposing a sparse mean-field theory. By revealing non-random correlations between intra-class variability and centroids, we demonstrate that classification is governed by a sparse centroid alignment structure, enabling an analytical prediction model dependent on only a few centroid coordinates. Integrating global radius renormalization with cross-modal geometric analysis, this framework accurately predicts classification accuracy across diverse architectures and elucidates systematic scaling laws of geometric quantities with respect to model size. Consequently, this work establishes a novel paradigm for understanding the geometric nature of universal cross-modal representations.
📝 Abstract
Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensional structure makes them difficult to interpret. We show that classification tasks give rise to a universal representational geometry, shared across state-of-the-art models in vision, audio, and language processing. The key structure is that within-class variability is not random in representation space. Instead, its classifier-relevant component has strong and structured correlations with the class's own centroid and with the centroids of its competing classes. Building on this observation, we derive an analytical mean-field theory governed mainly by the variability along true-class and rival-class centroid coordinates, together with a global renormalization of the class radius that compensates for the non-Gaussian statistics of real representations. The theory accurately predicts classification accuracy across architectures and modalities. The relevant geometric quantities improve systematically with model scale, mirroring the observed gains in accuracy. A striking feature of the theory is its sparsity: accurate prediction requires only a small set of centroid coordinates associated with the true class and its strongest rivals - connecting our framework to sparse-feature extraction approaches such as sparse autoencoders. Together, these results provide a parsimonious predictive theory of neural representations and suggest that classification in deep networks is governed by a sparse, centroid-aligned structure embedded within the full high-dimensional representation space.
Problem

Research questions and friction points this paper is trying to address.

Neural Representations
Interpretability
Representational Geometry
Cross-modal Classification
Predictive Theory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Prototype Code
Universal Representational Geometry
Analytical Mean-Field Theory
Cross-Modal Classification
Centroid-Aligned Structure
🔎 Similar Papers
No similar papers found.