M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark
Current text-to-image models exhibit limited capability in aligning generated images with complex textual descriptions involving multiple instances, diverse categories, and intricate semantic relationships; moreover, fine-grained evaluation benchmarks strongly correlated with human judgment remain scarce. To address these challenges, we introduce M3R-Bench—the first large-scale, multi-category, multi-instance, multi-relation image-text alignment benchmark—and propose Revise-Then-Enforce, a training-free post-editing method. We further design AlignScore, an automatic metric that jointly models object detection and semantic parsing to assess both visual entities and their relational structure, achieving strong correlation with human preferences (ρ > 0.85). Experiments reveal significant performance bottlenecks of mainstream open-source diffusion models on M3R-Bench, confirming its rigor; Revise-Then-Enforce consistently improves alignment quality across models—including Stable Diffusion—yielding an average AlignScore gain of 12.7%. This work establishes a new paradigm for evaluating and optimizing image-text alignment.