TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决流式输入下的视觉定位问题,TempoGround通过状态感知的跨帧对应和强化学习方法,实现准确且一致的视觉定位。
📝 Abstract
Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progress on single-frame and video-based visual grounding, yet under streaming inputs they still suffer from identity drift, cross-frame inconsistency, and fragile localization under partial occlusion. To address these issues, we present TempoGround, a VLM-native framework that detects cross-frame object correspondence and explicitly models object presence states, thereby enabling accurate and consistent visual grounding under streaming inputs. The key is a curriculum prediction mechanism guided by state-aware cross-frame correspondence: TempoGround resolves 2D instance association, predicts whether each object newly enters, continues in, or leaves the view, decodes the 2D box, and then lifts it to a camera-frame 3D box. As token-level supervision alone cannot capture the geometric objectives of streaming grounding, we further introduce Streaming Grounding Reinforcement (SGR), which optimizes TempoGround with verifiable Grounding, Identity, and Consistency rewards, jointly reinforcing persistent localization and temporally consistent predictions. We carefully design a three-stage training strategy and train TempoGround on large-scale data. We evaluate visual grounding under causally streaming inputs on multiple challenging benchmarks: TempoGround improves F1_2D@0.5 and F1_2D@0.95 by 4.4 and 0.5 on average, and F1_3D@0.25 and AP_3D by 6.2 and 7.5, respectively. These results demonstrate that TempoGround provides a practical foundation for visual grounding under streaming inputs.
Problem

Research questions and friction points this paper is trying to address.

Visual Grounding
Streaming Inputs
Identity Drift
Cross-Frame Inconsistency
Partial Occlusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

state-aware
cross-frame correspondence
streaming grounding reinforcement
visual-language models
temporal consistency
L
Leqian Ding
Xi’an Jiaotong University, Xi’an, China
J
Junning Qiu
EngineAI, Shenzhen, China
M
Manwen Yang
Xi’an Jiaotong University, Xi’an, China
Yu Guo
Yu Guo
Xi’an Jiaotong University
6D pose estimationtime series predictiongraph learning
F
Fei Wang
Xi’an Jiaotong University, Xi’an, China