A Low-Latency Interactive System for Real-Time Video Understanding Based on VLMs

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于边缘-云的实时视频理解系统,集成了多种视觉-语言模型后端,并通过WebRTC实现了低延迟交互。
📝 Abstract
Vision-language models are extending video understanding from offline clip analysis to continuous interactive streaming, but most research still emphasizes model capability rather than deployable low-latency interaction. This paper presents a unified edge-cloud system for real-time video VLM applications. Lightweight phone, smart glasses, PC, and pseudo-replay clients publish video and speech to a server runtime that provides shared ASR/TTS, session orchestration, backend adaptation, response delivery, and archive-backed measurement. The system integrates six representative video VLM backends with streaming or interaction-oriented capabilities and evaluates them across backend runtime, media transport, client-observed latency, and interaction behavior. With suitable backend selection and the WebRTC path, the tested system reaches approximately 0.9 to 1.0 s to first VLM text and 1.3 to 1.5 s to first non-silent TTS audio, while exposing backend adaptation costs and differences in real-time interaction behavior.
Problem

Research questions and friction points this paper is trying to address.

Low-Latency
Real-Time Video Understanding
Vision-Language Models
Interactive System
Innovation

Methods, ideas, or system contributions that make the work stand out.

Real-time Video Understanding
Edge-Cloud System
Low-Latitude Interaction
Visual-Language Models (VLMs)
WebRTC
🔎 Similar Papers
P
Punan Dai
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, China
J
Jun Xu
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, China
B
Bingcong Lu
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, China
Zhengxue Cheng
Zhengxue Cheng
Assistant Researcher, Shanghai Jiao Tong University
Video and Image CodingComputer VisionImage Quality Assessment
H
Hongwei Hu
Ant Group, Shanghai, China
R
Ronghua Wu
Ant Group, Shanghai, China
Li Song
Li Song
Professor of Electronic Engineering, Shanghai Jiao Tong University
Video CodingImage ProcessingComputer Vision