Few-Shot Video Recognition via Hierarchical Metric Learning

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对少样本视频识别问题,提出了一种层次度量学习方法HML-FSAR,通过增强空间模块和多层次度量策略来优化特征紧凑性和类别区分性。
📝 Abstract
Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail to sufficiently exploit rich cross-frame global spatial information in videos. Even existing multi-level metric schemes only impose parallel prototype constraints on intermediate layers, without progressive supervision along the full feature pipeline, which results in limited generalization ability of the learned class prototypes. Inspired by this, we present a novel method, hierarchical metric learning for few-shot action recognition (HML-FSAR). First, a spatial-enhanced module is developed to capture cross-frame global spatial representations. Combined with temporal MHA, heterogeneous alignment, spatial-temporal feature fusion and dictionary learning modules, it constructs the complete feature processing pipeline. Second, a hierarchical metric learning (HML) strategy is embedded into HML-FSAR. Composed of center metric, alignment metric, contrastive metric, dictionary metric and prototype metric, HML imposes progressive multi-stage complementary constraints from frame-level representations to final class prototypes, so as to jointly optimize feature compactness, heterogeneous spatial-temporal alignment, inter-class discriminability and anti-noise robustness. The proposed HML-FSAR method is validated on five widely-used FSAR datasets, and experimental results fully demonstrate its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

Few-shot action recognition
Cross-frame global spatial information
Hierarchical metric learning
Feature compactness
Inter-class discriminability
Innovation

Methods, ideas, or system contributions that make the work stand out.

hierarchical metric learning
spatial-enhanced module
temporal MHA
heterogeneous alignment
spatial-temporal feature fusion
🔎 Similar Papers
J
Jiaxin Zhang
School of Information Science and Engineering, University of Jinan, Shandong, China
H
Haoran Gao
School of Information Science and Engineering, University of Jinan, Shandong, China
X
Xizhan Gao
School of Information Science and Engineering, University of Jinan, Shandong, China
Zihao Dong
Zihao Dong
CS PhD Student, Northeastern University
VerificationRoboticsComputer Vision
T
Tingwei Wang
School of Information Science and Engineering, University of Jinan, Shandong, China
Sijie Niu
Sijie Niu
University of Jinan
Medical Image ComputingPattern Recognition