PRISM: Enhancing Protein Inverse Folding through Fine-Grained Retrieval on Structure-Sequence Multimodal Representations
Protein inverse folding—generating sequences compatible with a target 3D structure—is highly challenging due to the vast sequence space and complex local structural constraints. This paper proposes a fine-grained multimodal retrieval-augmented generation framework: first, retrieving evolutionarily conserved local structure–sequence motifs from a natural protein database; second, integrating retrieved motifs with the target structure via a hybrid self-cross-attention decoder; and third, explicitly modeling and reusing these motifs through a latent-variable probabilistic model. To our knowledge, this is the first approach to incorporate fine-grained, naturally occurring structure–sequence co-occurrence patterns into inverse folding. Evaluated on five standard benchmarks, our method achieves significant improvements in amino acid recovery rate and perplexity, while simultaneously enhancing foldability metrics—including RMSD, TM-score, and pLDDT—demonstrating both effectiveness and strong generalization capability.