Surgical Video Generation From Diffusion to World Models: A Survey

๐Ÿ“… 2026-08-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้€š่ฟ‡ๆ— ๆกไปถ็”Ÿๆˆใ€ๆœ‰ๆกไปถ็”Ÿๆˆๅ’Œไธ–็•Œๆจกๅž‹็”Ÿๆˆไธ‰็งๆ–นๆณ•่งฃๅ†ณๆ‰‹ๆœฏ่ง†้ข‘ๆ•ฐๆฎ็จ€็ผบ้—ฎ้ข˜๏ผŒๆ้ซ˜ๆ‰‹ๆœฏๆจกๆ‹ŸไธŽ่ฎญ็ปƒ็š„่ดจ้‡ใ€‚
๐Ÿ“ Abstract
Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.
Problem

Research questions and friction points this paper is trying to address.

Surgical Video Generation
Data Scarcity
Clinical Plausibility
Generalization
Physical Realism
Innovation

Methods, ideas, or system contributions that make the work stand out.

surgical video generation
causal dynamics modeling
data scarcity
๐Ÿ”Ž Similar Papers
No similar papers found.
Fuxiang Huang
Fuxiang Huang
The Hong Kong University of Science and Technology (HKUST)
Multimodal LearningFoundation model for Vertical DomainDomain Adaptation
Chenxu Zhang
Chenxu Zhang
ByteDance Inc.
Computer GraphicsComputer VisionAI
L
Liang Han
School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China
L
Lei Zhang
School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China