Replacing Labeled Real-image Datasets with Auto-generated Contours
Pretraining vision transformers (ViTs) typically relies on large-scale real-image datasets, raising concerns regarding data privacy, environmental cost, and annotation effort. Method: We propose Formula-Driven Supervised Learning (FDSL), the first framework enabling ViT pretraining exclusively on synthetically generated contour images—without real images, human annotations, or self-supervision. Contours are procedurally generated via mathematical formulas, yielding controllable complexity, zero bias, zero cost, and zero privacy risk. Contribution/Results: We demonstrate that contour structures alone encode sufficient semantic information for effective representation learning, and that moderately increasing pretraining task difficulty improves transfer performance. A ViT-Base pretrained via FDSL achieves 82.7% top-1 accuracy on ImageNet-1K fine-tuning—surpassing the ImageNet-21K baseline (81.8%). This establishes that purely synthetic contour data can match—or even exceed—the efficacy of large-scale real-image pretraining, opening a new pathway toward green, trustworthy, and interpretable vision foundation models.