Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment
This work addresses color drift, boundary ambiguity, and the limitation of latent-feature decoders treating images merely as intermediate visualizations in generative semantic segmentation. To overcome these issues, the authors propose the Semantic Prism framework, which employs diffusion distillation to construct a single-step generator, integrates a fixed class-color codebook to establish an explicit probabilistic interface, and introduces a hierarchical evidence alignment mechanism in logit space to predict residuals while preserving the image-defined interface as the reference for the final distribution. The study also innovatively presents C-IHD, an error-ranking method that requires no additional predictor. Experiments demonstrate that the approach achieves 72.07% mIoU (+11.39) with an ECE of 0.41% on Cityscapes, and attains 62.22% and 46.89% mIoU on BDD100K and ACDC, respectively, while significantly improving pixel-wise error ranking AUPR—e.g., from 0.6580 to 0.7557 on ACDC.