World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge
Text-to-image (T2I) models suffer significant performance degradation when generating novel or out-of-distribution (OOD) entities due to knowledge cutoff and insufficient semantic alignment. To address this, we propose World-To-Image—a framework that employs a web-search agent to dynamically retrieve relevant online images and integrates multimodal prompt optimization for real-time external knowledge injection and prompt enhancement. Our method requires no model fine-tuning and achieves knowledge-guided generation in only 2.7 iterations on average. On the NICE benchmark, it improves semantic accuracy by 8.1% over state-of-the-art approaches; it also achieves superior semantic consistency and visual aesthetic quality under LLMGrader and ImageReward evaluations. The core contribution is a scalable, low-overhead “retrieve–optimize” closed loop—marking the first integration of embodied agents into T2I generation to tackle OOD challenges.