๐ค AI Summary
This work addresses the challenges of semantic ambiguity and inaccurate transparency estimation in open-world image matting, primarily caused by foreground diversity and sparse boundaries. To tackle these issues, we propose RenderMatte, a novel framework that leverages full-parameter fine-tuning of FLUX.1 Kontext combined with image editing priors to achieve structure-preserving alpha prediction. We introduce a group-wise relative alpha alignment mechanism to enhance multi-sample consistencyโa first in the field. Additionally, we construct the first large-scale synthetic dataset featuring precise, hair-level alpha annotations, effectively overcoming the lack of high-quality boundary supervision in real-world data. Extensive experiments demonstrate that RenderMatte achieves state-of-the-art performance across all benchmarks, significantly improving both accuracy and robustness for high-fidelity image matting in open scenarios.
๐ Abstract
Image matting is an essential enabling technology for modern visual content production, where foreground extraction determines the realism and editability of downstream creation workflows. However, precise alpha estimation in open-world scenes remains challenging because real foregrounds exhibit highly diverse appearances and opacity patterns. This makes existing methods struggle with semantic ambiguity and fine-grained opacity variation, especially in sparse boundary regions that are fragile and difficult to supervise. To address this gap, we present RenderMatte, a trimap-guided matting framework that adapts FLUX.1 Kontext through full-parameter fine-tuning, leveraging image editing priors for structure-preserving alpha prediction. During supervised adaptation, an alpha-edge objective preserves the latent flow-matching signal while strengthening pixel-space boundary supervision. We further introduce group-relative alpha alignment for post-training. It compares multiple mattes sampled under the same trimap condition using matting-specific rewards for alpha accuracy, boundary fidelity, trimap compliance, and compositional consistency. To overcome the lack of precise edge annotations, we construct the RenderMatte dataset, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets. It features exact strand-level alpha annotations and diverse background composites. Experiments show state-of-the-art performance across all benchmarks, demonstrating a scalable path toward high-fidelity matting in open-world scenes.