Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决边缘设备处理长视频时计算和带宽受限的问题,提出了一种基于视觉需求路由的框架,通过单次字幕生成和按需帧检索来优化资源使用。
📝 Abstract
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We observe a visual-textual duality: language memories carry long-range temporal structure better than dense frames, while pixels remain decisive for attribute-level perception. Building on this insight, we propose Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework. The edge runs a single offline captioning pass that builds a dual-track narrative index, an event-level story skeleton plus a clip-level micro-log, cached and reused across queries without re-captioning. At query time, a cloud-side MLLM reasons over the index in a story-first loop centered on a lightweight Visual-Need Router: a per-query gating module that triggers bounded keyframe retrieval only for perceptual questions (appearance, on-screen text, attribute disambiguation) and keeps temporal-structural questions in language space. The router turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length. Experiments on long-video benchmarks demonstrate strong accuracy-efficiency trade-offs while substantially reducing online visual processing.
Problem

Research questions and friction points this paper is trying to address.

long-video understanding
edge devices
compute and bandwidth budgets
temporal structure
visual attributes
Innovation

Methods, ideas, or system contributions that make the work stand out.

budget-aware
edge-cloud agentic framework
visual-need router
dual-track narrative index
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Weitong Cai
Weitong Cai
Queen Mary University of London
H
Hang Zhang
Independent Researcher
Y
Yukai Huang
Durham University
Y
Yiqiao Xie
Imperial College London
S
Shan Gao
Huawei
Jiankang Deng
Jiankang Deng
Imperial College London
Computer VisionMachine Learning
S
Songcen Xu
Huawei
Jifei Song
Jifei Song
Huawei Noah’s Ark Lab
Neural RenderingComputer VisionDeep LearningImage ProcessingSpeech Processing
Z
Zhensong Zhang
Huawei