CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
This work addresses the single-image content-style decomposition (CSD) problem, aiming for high-fidelity content extraction and controllable style transfer. We propose a scale-aware disentanglement framework grounded in visual autoregressive modeling: (i) a scale-aware alternating optimization strategy to strengthen content-style separation; (ii) an SVD-driven style rectification module to suppress content leakage into style representations; and (iii) an enhanced key-value memory mechanism to improve identity consistency across stylized outputs. To further boost generalization, we introduce scale-aligned training and targeted data augmentation. Additionally, we release CSD-100—the first dedicated benchmark for CSD evaluation. Extensive experiments demonstrate that our method achieves significant improvements over state-of-the-art approaches in both content fidelity and style consistency, enabling more flexible and precise visual synthesis with enhanced controllability.