Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决视频语言建模中长时间动态表示的问题,本文通过引入具有细粒度时间标注的Kairos数据集来支持更精细和长范围的视频内容理解。
📝 Abstract
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. Kairos provides a general-purpose foundation for modeling visual experiences over time.
Problem

Research questions and friction points this paper is trying to address.

video-language modeling
temporal alignment
fine-grained annotation
visual dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

time-resolved annotations
fine-grained video-language modeling
long-duration videos
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13