🤖 AI Summary
该研究通过集成视频-AI平台,利用动作分割、物体检测与跟踪及大语言模型等技术,为显微外科手术训练提供实时反馈,解决专家评审难以规模化的问题。
📝 Abstract
Developing microanastomosis skill requires repeated practice with timely, action-specific feedback, yet expert review of lengthy microscope videos does not scale to frequent or distributed training. We present an integrated video-AI platform that turns a complete simulated procedure into inspectable, interactive feedback through three connected modules. First, a proposed transformer segments the video into six surgical actions. Second, object detection and tracking localize instrument tips within each action; the resulting kinematic features and action statistics drive supervised classification of five NOMAT-aligned performance dimensions. Third, a grounded large language model (LLM) uses these structured outputs to answer user questions about the current scene, actions, motion, and predicted performance through a unified interface. In a two-site study, 17 participants completed 72 procedures comprising 576 suture placements. The action-segmentation module achieved 87.66\% accuracy and 82.86\% F1, increasing to 93.62\% and 88.32\% after workflow-aware refinement. The five performance classifiers achieved 76.0\% mean accuracy, with Cohen's $\kappa$ from 0.63 to 0.93. Although the language interface and educational benefit require prospective evaluation, these results establish the technical basis for an expert-supervised platform that can shorten review, expose the evidence behind performance estimates, and support scalable formative microsurgical training.