VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
Existing video style transfer methods are often hindered by the scarcity of large-scale triplet data and effective modeling paradigms, leading to temporal inconsistency, fragile handling of occlusions, and flickering artifacts. To address these limitations, this work introduces VISTA-1000, the first large-scale synthetic dataset with aligned style, content, and motion, encompassing 1,000 distinct artistic styles. Building upon this dataset, we propose a context-aware transfer framework based on diffusion Transformers, augmented with a lightweight style adapter for robust style representation. By integrating joint modeling with a disentanglement strategy, our approach significantly outperforms existing methods in terms of style fidelity, temporal coherence, and content preservation, effectively suppressing flickering and drift artifacts.