Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入ProbeGen基准评估图像生成器在零样本设置下的视觉感知能力,比较了20种模型在11个基准测试中的表现。
📝 Abstract
Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting. We introduce ProbeGen, a benchmark for zero-shot generative perception that casts monocular depth estimation, referring/reasoning segmentation, and object counting as conditional generation tasks specified through text prompts, and compares 20 models in total---including proprietary and open-weight image generators, specialist perception models, and MLLMs---across 11 published benchmarks. We observe that pretrained image generators show measurable zero-shot perceptual competence, but with a clear trade-off: specialist models remain stronger for in-distribution accuracy and efficiency, while generative models are often more robust under distribution shift and better at compositional semantic reasoning. We hope this study helps establish zero-shot generative perception as a meaningful research direction and provides a useful foundation for future work at the intersection of visual generation and understanding.
Problem

Research questions and friction points this paper is trying to address.

zero-shot
image generators
visual perception
conditional generation
text prompts
Innovation

Methods, ideas, or system contributions that make the work stand out.

zero-shot generative perception
text prompts
conditional generation tasks
distribution shift
compositional semantic reasoning
🔎 Similar Papers
No similar papers found.