Compositional Feature Augmentation for Unbiased Scene Graph Generation
Scene Graph Generation (SGG) suffers from severe predicate long-tail distribution, where conventional re-sampling–based debiasing methods fail to improve tail-predicate performance—primarily due to insufficient modeling of relational triplet feature diversity. This paper introduces, for the first time, a feature disentanglement perspective: it decomposes triplet representations into intrinsic (predicate semantics) and extrinsic (contextual dependency) components. Building upon this, we propose a plug-and-play replace-mix augmentation strategy that enhances tail-predicate feature diversity without modifying the backbone architecture. The method is model-agnostic and computationally efficient. Evaluated on Visual Genome (VG) and PIC benchmarks, it achieves state-of-the-art performance across multiple metrics, notably boosting tail-predicate Recall@100 by a significant margin. Moreover, it seamlessly integrates with diverse SGG frameworks, demonstrating broad compatibility and practical utility.