Analyzing Visual Aircraft Representations with Sparse Autoencoders
This work addresses the limited interpretability of internal representations in vision models, which hinders understanding of their decision-making mechanisms. For the first time, sparse autoencoders are applied to decompose intermediate features of ConvNeXt trained on the FGVC-Aircraft dataset. Through analyses of activated image patches, activation magnitudes, and class selectivity, the study reveals that multiple sparse features correspond to identifiable aircraft parts or semantic visual patterns. Ablation studies linking features to input space, together with quantitative class selectivity metrics, systematically demonstrate these features’ influence on classification confidence and decision boundaries. The investigation also highlights inherent limitations, including feature ambiguity and coarse spatial localization, offering critical insights into the trade-offs between sparsity, interpretability, and representational fidelity in deep visual models.