Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Co-Annotator,通过专家提炼的视觉转换器和视觉-语言模型来指导视网膜OCT图像分析与文档记录,提高诊断效率和准确性。
📝 Abstract
Clinical AI often optimizes predictive performance without engaging how clinicians decide where to look and what to write. We present Co-Annotator, which distills expert gaze and dictation into two guidance components: a gaze-aligned Vision Transformer producing fixation-aligned areas of interest (AOIs), and an ontology-bounded vision-language model (VLM) that pre-fills editable biomarker summaries for retinal optical coherence tomography (OCT). We first collect expert gaze and dictations (US1) to train the models, significantly improving diagnostic accuracy and biomarker generation. We then deploy the system with ophthalmology residents: a controlled resident study (US2) confirmed each modality is safe and independently beneficial, with AOI guidance producing lasting perceptual efficiency gains through post-guidance carryover and VLM guidance more than doubling biomarker documentation breadth. In a combined deployment across two academic institutions (US3), providing both modalities simultaneously produced efficiency gains that substantially exceeded either modality alone: correct diagnoses per minute increased by 40% and comment editing time fell by 67%, without compromising diagnostic accuracy. Notably, neither modality improved efficiency during guidance in US2, which makes the in-guidance efficiency gain under combined guidance in US3 the more striking result. Expert-distilled multimodal guidance can remove two distinct clinical workflow bottlenecks at once (visual search overhead and documentation burden) without compromising the diagnostic accuracy clinicians already achieve.
Problem

Research questions and friction points this paper is trying to address.

Clinical AI
Age-Related Macular Degeneration
Visual and Documentation Guidance
Workflow Bottlenecks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Transformer
vision-language model
multimodal guidance
clinical workflow efficiency
Z
Ziheng "Leo" Li
Columbia University
B
Benjamin Freeman
Columbia University
A
Akshay Raman
Columbia University
K
Kavin Aravindhan Rajkumar
Columbia University
X
Xinxin Fang
Columbia University
R
Rishabh Srivastava
Columbia University
Steven Feiner
Steven Feiner
Professor of Computer Science, Columbia University
Human-Computer InteractionAugmented RealityVirtual Reality3D User InterfacesWearable Computing
Kaveri A. Thakoor
Kaveri A. Thakoor
Columbia University