ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ViTAMINS方法,通过在自监督视觉Transformer预训练中引入合成难负样本以提升表示质量,并在多个任务上展示了其优越性能。
📝 Abstract
We introduce ViTAMINS, a method that integrates synthetic hard negatives into unsupervised vision transformer pretraining to improve representation quality. Our approach is thoroughly benchmarked on ImageNet and transfer learning, image retrieval, copy detection, and image, video segmentation tasks. Notably, our proposed negatives give rise to emergent properties, where learned representations contain explicit information about the semantic content of an image and serve as excellent classifiers (up to +11.3% over baselines). ViTAMINS achieves these benefits through simple modifications to existing contrastive frameworks and outperforms competing methods while being more resource efficient, e.g., our ViT-B surpasses V-JEPA with ViT-L. Our findings motivate reconsidering contrastive learning as a simpler yet powerful alternative to dominant generative and self-distillation approaches.
Problem

Research questions and friction points this paper is trying to address.

self-supervised
vision transformers
synthetic hard negatives
representation quality
contrastive learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic hard negatives
contrastive learning
representation quality
unsupervised vision transformer pretraining
🔎 Similar Papers
💼 Related Jobs
No related jobs found.