Role-SynthCLIP: A Role Play Driven Diverse Synthetic Data Approach
Existing synthetic data generation methods primarily focus on scaling data volume, yet suffer from limited semantic diversity and coarse-grained image-text alignment, leading to redundancy and superficial descriptions. To address this, we propose Role-SynthCLIP, a novel framework that introduces a role-playing mechanism—e.g., “Composition Analyst” and “Context Interpreter”—to guide multimodal large language models (MLLMs) in generating semantically rich, fine-grained aligned image-text pairs from diverse perspectives, without increasing dataset size. This enhances both semantic diversity and cross-modal alignment precision. The high-quality synthetic data produced by Role-SynthCLIP is employed for contrastive pretraining of CLIP. Using only 1 million synthetic pairs, our method achieves 64.1% Recall@1 on MS COCO—outperforming the 5-million-pair baseline by 2.8 percentage points. These results empirically validate the efficacy of role-driven prompting in improving synthetic data quality and downstream representation learning.