PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding
Existing sparse autoencoders struggle to effectively interpret pairwise representations in Pairformer-like protein co-folding models, often leading to feature explosion and an inability to model the joint distribution of sequence and pairwise features. This work proposes PairSAE, a novel framework that, for the first time, employs N-mode SVD to compress pairwise tensors into token-centric interaction roles and introduces a shared sparse autoencoder to jointly reconstruct both sequence and pairwise representations. By circumventing the quadratic growth inherent in conventional sparse autoencoders along pairwise dimensions, PairSAE extracts highly interpretable features on the PLINDER complex that align closely with UniProt functional annotations and accurately predicts Boltz-2 binding affinities, thereby uncovering structurally meaningful biological concepts learned internally by the model.