A Deeper Analysis of Block-Sparse Featurizers

๐Ÿ“… 2026-08-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๅˆ†ๆžไบ†ๅ—็จ€็–็‰นๅพๆๅ–ๅ™จ๏ผˆBSF๏ผ‰็š„ไผ˜็ผบ็‚น๏ผŒ้’ˆๅฏนๅ…ถๅญ˜ๅœจ็š„็‰นๅพๅˆ†่ฃ‚็ญ‰้—ฎ้ข˜๏ผŒๆๅ‡บไบ†ไธ€็ง็ซž่ต›Top-K้€‰ๆ‹ฉ่ง„ๅˆ™็ญ‰ๆ”น่ฟ›ๆ–นๆณ•ใ€‚
๐Ÿ“ Abstract
The recently introduced block-sparse featurizer (BSF; Fel et al., 2026) is similar to a sparse autoencoder (SAE), but its atomic unit is a small subspace (a block of directions) rather than a single direction. It is designed for features that live on low-dimensional manifolds, which are especially frequent in vision. This work studies the BSF's strengths and weaknesses, finding how it still somewhat suffers from classic SAE failure modes, like feature splitting and composition. We propose several architectural changes to the BSF, including a Tournament Top-K selection rule that significantly reduces feature splitting, and we also extend the block paradigm to the crosscoder.
Problem

Research questions and friction points this paper is trying to address.

Block-Sparse Featurizer
Low-Dimensional Manifolds
Feature Splitting
Feature Composition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Block-Sparse Featurizer
Tournament Top-K selection rule
crosscoder
๐Ÿ”Ž Similar Papers
A
Alexandru-Iulius Jerpelea
Columbia University
Amith Ananthram
Amith Ananthram
Columbia University
NLPCVAI