Boosting Data Augmentation with Stochastic Weight Averaging

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational costs of data augmentation ensembles and their inefficiency in leveraging task symmetries by proposing Stochastic Weight Averaging (SWA) as a replacement for repetitive ensembling. Through approximation analysis via the Ornstein-Uhlenbeck process, we reveal that SWA enhances model equivariance beyond conventional performance gains in the infinite-width limit. Experiments on visual and graph classification tasks demonstrate the method’s superiority across both discrete and continuous symmetries. These findings validate SWA as an effective alternative to traditional ensembling, providing new theoretical foundations and a practical paradigm for efficiently exploiting data augmentation. This work thus bridges the gap between computational efficiency and symmetry-aware learning, offering significant implications for scalable representation learning in structured domains.
📝 Abstract
The symmetries of a learning task have become an important factor in designing modern deep learning solutions. Data augmentation is a straightforward and effective way of incorporating symmetries into a generic neural network. Recent results show that infinitely large deep ensembles show perfect symmetry when trained on augmented data. However, since training ensembles requires repeating the training process many times, this method is costly. In this work, we study stochastic weight averaging (SWA) as an alternative ensembling technique that does not require repeated training runs. We analyze SWA by approximating the stochastic training trajectory at the end of training with an Ornstein--Uhlenbeck process. We show that in the infinite-width limit, SWA on augmented data provides an equiviariance boost that goes beyond what could be expected from the performance increase due to SWA alone. We verify our results with extensive numerical experiments on numerous models spanning computer vision and graph classification with both discrete and continuous symmetries.
Problem

Research questions and friction points this paper is trying to address.

Data Augmentation
Stochastic Weight Averaging
Symmetry
Deep Ensembles
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Weight Averaging
Data Augmentation
Equivariance
Ornstein-Uhlenbeck Process
Symmetry
💼 Related Jobs
No related jobs found.
L
Longde Huang
Department of Mathematical Sciences, Chalmers University of Technology and the University of Gothenburg, SE-412 96 Gothenburg, Sweden
Axel Flinth
Axel Flinth
Assistant professor, Umeå University
Compressed sensingsignal processingdata sciencegeometric deep learning
J
Jan E. Gerken
Department of Mathematical Sciences, Chalmers University of Technology and the University of Gothenburg, SE-412 96 Gothenburg, Sweden