Style-CCL: Content-Preserving Style Transfer via Curriculum Continual Learning
Existing diffusion Transformers struggle to effectively disentangle content and style in style transfer, often being dominated by semantic-level style cues, which leads to insufficient texture learning and content distortion. To address this, this work proposes Style-CCL—the first multi-stage training framework that integrates curriculum learning with continual learning. It trains a dual-branch SC-DiT model following a structured progression from “semantic to texture” and “clean to synthetic” data, while incorporating random memory replay to mitigate catastrophic forgetting. The approach leverages independent RoPE embeddings, causal masking, and reverse triplets to construct a million-scale dataset. Evaluated on style similarity, content consistency, and aesthetic quality, Style-CCL achieves state-of-the-art performance, significantly enhancing both stylistic expressiveness and content fidelity.