MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of weak temporal reasoning, detail loss, and poor transferability in long-duration surgical video understanding by proposing a perception-reasoning decoupled agent framework. The architecture employs a text orchestrator to plan evidence collection alongside frozen visual sub-agents for tool execution, integrated with a gradient-free heuristic skill distillation mechanism that adaptively evolves a reusable external skill library from low-scoring trajectories. Requiring only approximately one hundred annotated samples for skill retrieval optimization, this approach comprehensively outperforms existing vision-language models and video agents on both proprietary neurosurgical and public benchmarks. Consequently, the proposed method significantly enhances out-of-domain generalization capabilities and temporal reasoning accuracy within long-video contexts, offering a robust solution for complex surgical analysis tasks.
📝 Abstract
Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time. Existing approaches handle this poorly: a one-shot vision-language model (VLM) compresses the whole procedure to fit its context window and loses the detail a "before" or "after" question depends on, while video agents that train the model where to look are data-hungry and transfer poorly to out-of-domain surgery. We build an agent harness that separates reasoning from perception and improves by evolving context rather than optimizing weights. A text-only orchestrator plans which evidence to gather and issues an auditable sequence of tool calls, while frozen vision-language sub-agents execute each call over the pixels, viewing, cropping, inspecting frames, and retrieving external knowledge. We further propose a gradient-free, reward-gated Heuristic Skill Distillation loop that mines the agent's own low-scoring traces and keeps a candidate skill only when it raises a validation reward, yielding reusable retrieval skills, notably directed re-look. Growing an external skill library rather than tuning weights, the loop adapts from only about 100 labeled examples, far fewer than supervised or reinforcement fine-tuning requires. To evaluate this agent, we introduce MedClawBench, a de-leaked, doctor-grounded benchmark of 1,123 questions over self-built long neurosurgery recordings and a held-out public lecture-video test split. Across both datasets and all four evaluation dimensions, our agent consistently outperforms one-shot VLMs and general video-agent frameworks, with the largest gains on the long, out-of-domain neurosurgery videos. Project page: https://fyycs.github.io/medclaw/.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon surgical video reasoning
Temporal reasoning
Vision-language models
Out-of-domain generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heuristic Skill Distillation
Agent Harness
Long-Horizon Surgical Video Reasoning
Gradient-Free Adaptation
MedClawBench
Yingying Fan
Yingying Fan
School of Automation and Intelligence, Beijing Jiaotong University, Beijing 100044, China; Department of Electrical and Computer Engineering, University of Maryland, College Park, MD 20742, USA
Penghui Du
Penghui Du
Southern University of Science and Technology, Undergraduate
NeuroscienceMachine LearningfMRI imaging
Leyan Zhu
Leyan Zhu
UniPat.ai
Runze He
Runze He
Institute of Information Engineering, Chinese Academy of Sciences
Computer Vision
Z
Zimeng Wu
UniPat.ai
Y
Yuxuan Zhang
University of British Columbia
L
Liang Chen
UniPat.ai
J
Jiahao Xie
UniPat.ai
J
Jiangtang Wang
Suzhou Institute for Advanced Research, University of Science and Technology of China
Shuai Shao
Shuai Shao
University of Science and Technology of China
Complexity TheoryInformation Theory
A
Anchao Yang
Department of Neurosurgery, Beijing Tiantan Hospital, Capital Medical University
Yutong Bai
Yutong Bai
Postdoc, UC Berkeley
Artificial IntelligenceComputer VisionDeep Learning
Y
Yan Wang
School of Automation and Intelligence, Beijing Jiaotong University, No.3 Shangyuancun, Haidian District, Beijing, 100044, China