GraViT: Transfer Learning with Vision Transformers and MLP-Mixer for Strong Gravitational Lens Discovery
Automated strong gravitational lens detection in the massive image datasets expected from next-decade surveys (e.g., LSST) remains challenging due to limited labeled examples and high annotation costs. Method: We propose a transfer learning framework leveraging Vision Transformers and MLP-Mixers—two state-of-the-art vision architectures—and systematically evaluate their adaptability to lens classification. Our approach integrates multi-source data augmentation, fine-grained fine-tuning, model ensembling, and optimizer hyperparameter tuning to enhance few-shot generalization. Results: Fine-tuned on HOLISMOKES VI and SuGOHI X, ten transformer- and mixer-based models consistently outperform conventional CNN baselines in classification accuracy. Crucially, they achieve favorable trade-offs between inference speed and model complexity, demonstrating scalability and robustness for large-scale survey operations. This work establishes a new, extensible, and highly robust paradigm for automated lens detection—directly supporting dark matter mapping and precision cosmological parameter inference.