Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
Generative vision-language models (e.g., Stable Diffusion) excel in creative image synthesis but remain vulnerable to adversarial prompts that elicit unsafe, offensive, or culturally inappropriate outputs; existing defenses often compromise image quality or incur substantial computational overhead. This paper introduces VALOR—a modular, zero-shot proxy framework that enhances safety and utility in text-to-image generation via hierarchical prompt analysis and value-aligned reasoning. Its core contributions include: (i) integrated multi-level NSFW detection, cultural-value alignment, and intent disambiguation; (ii) LLM-driven selective prompt rewriting and optional stylistic regeneration; and (iii) semantics-preserving safe regeneration with dynamic role-instruction adaptation. Experiments demonstrate that VALOR achieves up to 100% suppression of unsafe outputs across diverse adversarial and culturally sensitive prompts, while preserving prompt fidelity, creativity, and functional utility.