Seeing What You Say: Expressive Image Generation from Speech
This work addresses end-to-end speech-to-image generation—producing semantically accurate and emotionally consistent images directly from raw speech, bypassing automatic speech recognition (ASR) as an intermediate step. To this end, we propose VoxStudio, a unified framework featuring: (i) a speech information bottleneck module that compresses raw speech into compact tokens encoding semantic, prosodic, and affective information; (ii) VoxEmoset, the first large-scale emotional speech–image paired dataset; and (iii) a cross-modal alignment mechanism enabling joint optimization of speech representations and the image generator. Experiments on SpokenCOCO, Flickr8kAudio, and VoxEmoset demonstrate substantial improvements in emotional consistency and visual fidelity of generated images. Results validate the critical role of paralinguistic modeling—particularly prosody and emotion—in speech-driven multimodal generation, establishing a novel paradigm for direct speech-conditioned visual synthesis.