EndoNav: Semantic-to-Geometric Grounding for Language-Guided Robotic Endoscopic Examination

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
EndoNav通过将高级外科医生命令转化为患者特定解剖结构中的自主内窥镜可视化行为,解决了机器人内窥镜在微创手术中缺乏有效视觉辅助的问题。
📝 Abstract
Minimally invasive procedures performed within confined anatomical spaces depend on continuous endoscopic visualization. Current robotic endoscope systems can stabilize or reposition an endoscope, but they do not possess relevant context to provide effective visualization assistance. We present EndoNav, an anatomy-grounded natural-language framework that translates high-level surgeon commands into autonomous endoscopic visualization behaviors within patient-specific sinonasal anatomy. Spoken surgeon commands are transcribed and interpreted by an endoscopic viewpoint agent conditioned on a patient-specific anatomical scene representation. Rather than generating robot motion directly, the viewpoint agent generates structured visualization objectives that are converted into target viewpoints and inspection trajectories, which are then executed through geometry-constrained endoscope motion planning and joint-space control. We evaluate EndoNav using a structured three-pass sinus examination across three CT-derived anatomical models. For one cadaveric specimen, autonomous visualization is compared with sinus examinations performed by two resident surgeons. EndoNav achieved mean visualization IoUs of 87.04% and 84.37% relative to the two surgeon examinations, compared with an inter-surgeon IoU of 87.44%, while recovering 92.91% and 93.20% of surgeon-observed anatomical surfaces, respectively. These results demonstrate the feasibility of grounding high-level anatomical commands into patient-specific geometric objectives and translating them into anatomically constrained robotic visualization behaviors.
Problem

Research questions and friction points this paper is trying to address.

robotic endoscope
minimally invasive procedures
visualization assistance
anatomical context
Innovation

Methods, ideas, or system contributions that make the work stand out.

anatomy-grounded natural-language framework
autonomous endoscopic visualization
structured visualization objectives
geometry-constrained motion planning
patient-specific anatomical models