Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决现有模型无法处理动态感知和预测的问题,本文提出了Human-JEPA模型,通过锚定预测方法在视频上进行训练,从而实现对人类行为的理解。
📝 Abstract
Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification collapse. Under frozen probes, Human-JEPA leads the pixel-anchored specialists on pose and person re-identification at 2.7 times fewer parameters, conceding high-resolution dense parsing, and its released predictor head is the first that does not degrade anticipation. A single safely adapted model thus serves both halves of understanding humans.
Problem

Research questions and friction points this paper is trying to address.

human-centric vision model
perception
anticipation
motion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-JEPA
anchored forecasting
dense perception
past-to-future split
anticipation
🔎 Similar Papers
No similar papers found.