AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the growing demand in digital applications for high-quality, semantically aligned, and contextually relevant AI-generated sound effects that exhibit both diversity and controllability. It presents a systematic review of sound effect generation models from the past five years, encompassing approaches driven by text, visual, audio, and multimodal inputs. By analyzing 30 peer-reviewed papers sourced from Google Scholar, IEEE Xplore, and ACM Digital Library, this work offers the first comprehensive comparison of how different input modalities influence generation performance. The review highlights advances in audio fidelity, semantic alignment, and temporal coherence, while identifying persistent challenges such as temporal synchronization and perceptual consistency. Furthermore, it underscores a significant gap between current objective evaluation metrics and human perceptual judgments.
📝 Abstract
Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momentum, this chapter reviews and analyzes recent AI-based generative models for sound effect synthesis, with a focus on how different input modalities (text, visual, audio, and multimodal) affect the quality, controllability, and contextual relevance of the generated audio. It examines 30 peer-reviewed articles sourced from Google Scholar, IEEE Xplore, and the ACM Digital Library, exploring the evolution of AI generative models over the past five years. The results show that multiple models achieved state-of-the-art performance, producing high-fidelity, semantically aligned, and increasingly temporally coherent sound effects across tasks. However, despite these advances, the review identifies persistent challenges, including limitations in temporal synchronization for complex multi-event scenarios, gaps between objective metrics and human perception, and trade-offs between controllability and generative diversity. Overall, the chapter highlights that AI-driven sound effect generation is progressing toward more adaptive, scalable, and context-aware systems, offering significant implications for future sound design workflows and interactive media applications.
Problem

Research questions and friction points this paper is trying to address.

sound effect generation
temporal synchronization
human perception
controllability
generative diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-based sound generation
input modalities
generative models
sound effect synthesis
multimodal audio generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sandy Abdo
Ontario Tech University, Oshawa, ON, Canada
Bill Kapralos
Bill Kapralos
Ontario Tech University
Immersive technologiesserious gamingvirtual simulationspatial soundmultimodal interactions
P
Priyamvada Tripathi
Durham College, Oshawa, ON, Canada
K
KC Collins
Carleton University, Ottawa, ON, Canada
A
Adam Dubrowski
Ontario Tech University, Oshawa, ON, Canada