RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出RefineRank,通过结合视觉语言模型和开放集检测器的优势,解决手术时空定位中理解问题与精确定位之间的矛盾。
📝 Abstract
Surgical spatio-temporal grounding (STG) requires locating, at each video time specified by a procedural question, the object that the question asks about. Existing approaches face a trade-off: vision language models understand the question context but produce imprecise coordinates, whereas open-set detectors provide localized candidate boxes whose confidence does not reflect which box answers the question. We introduce RefineRank, which closes this gap at the candidate-box level. A compact trainable module, RefineNet, combines the language and regional features of a frozen medical vision language model with the proposals of a frozen open-set detector: it predicts a bounded coordinate correction and a quality score for every candidate box, and a fixed decoding rule returns the original or refined box with the highest score. On the MedVidBench Official Rankings (Verified), RefineRank records 0.421 STG mIoU, the highest displayed STG score, while its global multi-metric rank is 11. In a controlled evaluation on separate training and evaluation videos, coordinate correction raises the candidate oracle upper bound from 0.6772 to 0.7302, and ranking the joint pool of original and refined candidates by their RefineNet scores improves STG mIoU from 0.2719 to 0.4534, whereas separately trained selectors over the same pool reach at most 0.4186. These results show that a small box-level module can reconcile question understanding with precise localization without retraining either backbone. Code is available at [https://github.com/linzhe001/RefineRank](https://github.com/linzhe001/RefineRank).
Problem

Research questions and friction points this paper is trying to address.

Surgical Spatio-Temporal Grounding
Vision Language Models
Open-Set Detectors
Innovation

Methods, ideas, or system contributions that make the work stand out.

RefineRank
Surgical Spatio-Temporal Grounding
RefineNet
Coordinate Correction
Quality Score
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
L
Linzhe Jiang
UCL Hawkes Institute, University College London, London, UK
J
Jiayuan Huang
Visual Understanding Research Group, Department of Informatics, King’s College London, London, UK
C
Changhao Zhang
UCL Hawkes Institute, University College London, London, UK
Chunyang Jiang
Chunyang Jiang
HKUST
Artificial IntelligenceNatural Language Processing
Z
Zhehua Mao
UCL Hawkes Institute, University College London, London, UK
M
Mobarak I. Hoque
Division of Informatics, Imaging and Data Sciences, University of Manchester, Manchester, UK