Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视频时空间定位中复杂查询问题,提出使用合成数据引擎生成难度分级的数据,并引入课程强化学习框架CurrSTVG提高模型处理复杂查询的能力。
📝 Abstract
Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely overlook the compositional queries prevalent in real-world scenarios, where a target must be disambiguated by jointly reasoning about its attributes and relations to other entities. To bridge this gap, we propose Compositional Spatio-Temporal Video Grounding (CompSTVG), a task that requires models to process complex textual queries where every intertwined attribute and relational cue is essential for disambiguation. To facilitate this task at scale, we build a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training. Built on this engine, we introduce STVG-CompBench, a benchmark stratified by explicit difficulty levels that jointly capture temporal complexity and spatial interference. Evaluating 11 representative STVG models on STVG-CompBench reveals that current models perform poorly on compositional queries, exhibiting a sharp performance drop that is typically obscured by overall dataset-level averages. We further construct synthetic training data and propose CurrSTVG, a curriculum reinforcement learning framework that delivers consistent gains, with the largest improvements observed on the most challenging compositional queries.
Problem

Research questions and friction points this paper is trying to address.

Spatio-Temporal Video Grounding
Compositional Queries
Attribute and Relational Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Spatio-Temporal Video Grounding
Synthetic Data Engine
Constraint Programming
STVG-CompBench
Curriculum Reinforcement Learning
🔎 Similar Papers
No similar papers found.
X
Xingjian Wang
Monash University
S
Shijian Wang
Monash University
Y
Yibo Wang
Monash University
Zihao Yu
Zihao Yu
University of Science and Technology of China
R
Runhao Fu
Monash University
Xuelian Cheng
Xuelian Cheng
Monash University
3D VisionMedical ImagingMachine Learning
Z
Zongyuan Ge
Monash University