GenCAMO: Scene-Graph Contextual Decoupling for Environment-aware and Mask-free Camouflage Image-Dense Annotation Generation
This work addresses the scarcity of high-quality, large-scale annotated camouflage imagery that hinders dense prediction tasks such as camouflaged object detection and open-vocabulary segmentation. To overcome this limitation, we propose an environment-aware, mask-free generative framework that leverages scene graph context disentanglement to jointly synthesize realistic multimodal camouflaged images along with their dense annotations—including depth maps, attribute descriptions, and textual prompts—thereby constructing GenCAMO-DB, the first large-scale synthetic dataset for this domain. Experimental results demonstrate that models trained on our synthesized data achieve significantly improved performance in complex camouflaged scenarios, validating both the effectiveness and generalization capability of the generated data.