SmolVLM: Redefining small and efficient multimodal models
To address the high GPU memory consumption and low inference efficiency of large vision-language models (VLMs) on mobile and edge devices, this work proposes an architecture–tokenization–data co-optimization paradigm tailored for resource-constrained scenarios. Methodologically, we design a lightweight Transformer backbone, introduce sparse image/video tokenization strategies, and construct a high-quality, compact multimodal dataset trained via curriculum learning. Our key contributions are: (1) SmolVLM-256M achieves <1 GB GPU memory usage during inference while outperforming Idefics-80B in accuracy; (2) the 2.2B-parameter variant attains state-of-the-art performance on both image and video understanding tasks with significantly lower memory footprint; and (3) this is the first demonstration of a small-parameter VLM systematically surpassing ultra-large models across multimodal understanding benchmarks—establishing a new paradigm for efficient VLM deployment on edge devices.