TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决零样本时序动作检测中语义区分度不足的问题,提出TF-CADE方法,通过文本-前景集中对齐增强文本与视频特征的一致性。
📝 Abstract
Zero-Shot Temporal Action Detection (ZSTAD) aims to lo- calize and recognize action instances from unseen action categories in untrimmed videos. Although existing meth- ods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing se- mantic distinctions between action classes, resulting in text- irrelevant predictions. To address this issue, we propose a Text-Foreground Concentrated Alignment for zero-shot temporal action DEtector (TF-CADE) that explicitly aligns textual information with action-relevant foreground regions. Specifically, we introduce Action Concentrate Aggregation (ACA), which extracts action concentrate scores to aggregate temporally informative video segments into a foreground- weighted video embedding. This foreground concentrated alignment enhances the semantic consistency between text and video features and improves inter-class discriminabil- ity. In addition, a Certainty-based Confidence Re-weighting (CCR) strategy refines per-snippet confidence scores by lever- aging foreground-aware similarity, effectively suppressing irrelevant action classes during inference. Extensive evalua- tions show that our TF-CADE not only achieves state-of-the- art performance under in-distribution settings but also excels in cross-dataset generalization to unseen action classes.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Temporal Action Detection
text-video alignment
semantic distinctions
action classes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-Foreground Concentrated Alignment
Action Concentrate Aggregation (ACA)
Certainty-based Confidence Re-weighting (CCR)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yearang Lee
Dept. of Artificial Intelligence, Korea University, Seoul, Korea
Ho-Joong Kim
Ho-Joong Kim
Korea University
computer vision
S
Seong-Whan Lee
Dept. of Artificial Intelligence, Korea University, Seoul, Korea