Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种JEPA风格框架,通过预测被遮罩的时空点管来解决4D点云视频自监督学习中注释成本高和基于重建预训练过分强调几何细节的问题。
📝 Abstract
Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining can overemphasize low-level geometric details. We propose a JEPA-style framework that learns from unlabeled spatiotemporal point clouds through latent point-tube prediction. Instead of reconstructing raw coordinates, the model masks spatiotemporal regions and predicts their target representations from visible context representations in feature space. To stabilize latent prediction, we incorporate Sketched Isotropic Gaussian Regularization, which encourages non-collapsed embeddings without relying on explicit reconstruction targets. This formulation aims to capture both spatial structure and temporal dynamics while keeping the pretraining objective aligned with downstream semantic recognition. Experiments on action and gesture recognition benchmarks show that the learned representations improve downstream fine-tuning, limited-label learning, and cross-dataset transfer. These results suggest that JEPA-style latent prediction is a promising alternative to reconstruction-centered pretraining for 4D point cloud videos.
Problem

Research questions and friction points this paper is trying to address.

self-supervised learning
4D point cloud videos
unlabeled spatiotemporal point clouds
latent prediction
reconstruction-based pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

JEPA-style framework
latent point-tube prediction
Sketched Isotropic Gaussian Regularization
4D point cloud videos