From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of domain expertise and adaptation challenges in frozen general-purpose vision-language models for endoscopic polyp reporting. We propose a context fusion framework that leverages implicit instruction and explicit transduction mechanisms, integrating retrieval-augmented generation with continuous expert token learning to achieve lightweight specialization without modifying model weights. Experiments demonstrate that this approach significantly outperforms baseline models while introducing only 0.006% additional parameters. Furthermore, when provided with accurate retrieved evidence, the method corrects 70.5% of errors associated with weight-based adaptation, effectively balancing a unified interface with the preservation of pretrained capabilities.
📝 Abstract
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.
Problem

Research questions and friction points this paper is trying to address.

Endoscopic Polyp Reporting
Vision-Language Models
Specialist Adaptation
Frozen VLM
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-Fusion Framework
Frozen VLM
Retrieval-Augmented Generation
Continuous Specialist Tokens
Endoscopic Polyp Reporting
🔎 Similar Papers
No similar papers found.
R
Ruijie Yang
Zhejiang University, Hangzhou, China; Shanghai Institute for Advanced Study, Zhejiang University, Shanghai, China; Shanghai Key Laboratory of MICCAI, Shanghai, China; Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, China
Y
Yan Zhu
Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China; Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China
P
Peiyao Fu
Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China; Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China
S
Siyuan Li
Zhejiang University, Hangzhou, China; Shanghai Key Laboratory of MICCAI, Shanghai, China; Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, China
T
Te Luo
Shanghai Key Laboratory of MICCAI, Shanghai, China; Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, China
Zhihua Wang
Zhihua Wang
City University of Hong Kong
Computer VisionBiomedical EngineeringRobotics
Q
Quanlin Li
Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China; Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China
P
Pinghong Zhou
Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China; Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China
Xian Yang
Xian Yang
University of Manchester
Artificial IntelligenceMachine LearningHealthcare AINatural Language Processing
Shuo Wang
Shuo Wang
Fudan University
AI for Multi-Modal MedicineMedical Image AnalysisBiomechanics