Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost and lack of spatial traceability in surgical multimodal large language models caused by dense visual tokens. The authors propose Slot2Text, a novel approach that replaces dense visual representations with a small number of region-labeled slots generated via self-supervised feature clustering, significantly reducing token count while enabling explicit region localization. Slot2Text employs a slot encoding mechanism and a dual-mode architecture—Slot2Text-Fast and Slot2Text-Reason—augmented with a region-mask association technique. Experiments demonstrate that Slot2Text-Fast maintains state-of-the-art performance while reducing total token consumption by 91.8% (visual prefix tokens decrease from 1,295 to 47), and Slot2Text-Reason achieves explicit alignment among region identity, spatial location, and language output, thereby enhancing traceable reasoning capabilities.
📝 Abstract
Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leading to costly inference and limited spatial traceability for generated answers. We present Slot2Text, a dual-mode surgical MLLM that replaces dense representations of visual input with a compact set of regions encoded as slot latents. Instead of relying on contrastive alignment of the visual encoder with language, Slot2Text groups self-supervised vision features into a few regions--slots that are consumed by the language model as area-labeled visual tokens. Slot2Text-Fast uses the slot prefix to answer surgical questions. Slot2Text-Reason also identifies and locates areas relevant for reasoning, linking language outputs to corresponding slot tokens, masks or regions. Experiments on multiple visual question answering and visual grounding benchmarks show that Slot2Text-Fast is competitive with state-of-the-art baseline at a much lower cost, reducing the average total token consumption by a 91.8\% and the visual prefix from 1,295 to 47 tokens (a 96.4\% reduction). Slot2Text-Reason trades additional tokens and latency for explicit area identities, locations, and traceable spatial evidence. These results establish compact slot latents as an efficient default visual interface for surgical MLLMs, with grounded reasoning invoked when greater spatial traceability is required.
Problem

Research questions and friction points this paper is trying to address.

surgical MLLMs
visual tokenization
spatial traceability
inference efficiency
dense visual tokens
Innovation

Methods, ideas, or system contributions that make the work stand out.

slot latents
visual tokenization
surgical MLLMs
spatial traceability
efficient inference
Guiqiu Liao
Guiqiu Liao
University of Pennsylvania
Surgical roboticsComputer visionMachine learning
Matjaž Jogan
Matjaž Jogan
University of Pennsylvania, University of Ljubljana
artificial intelligencecomputer assisted surgeryperception
D
Daniel A. Hashimoto
GRASP Laboratory, University of Pennsylvania; PCASO Laboratory, Department of Surgery, University of Pennsylvania; Department of Computer and Information Science, University of Pennsylvania