Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过将JEPA模型应用于点云数据,解决了几何观测下的潜在空间规划问题,采用三种JEPA设计并证明了在控制任务中的有效性。
📝 Abstract
JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction survives geometric observations is unclear: point clouds are sparse, unordered, and self-occluded, and with 0.3-15% of scene points moving, the slow-feature optimum of latent prediction compounds with the geometric shortcut of 3D self-supervision. We lift three canonical JEPA designs to point clouds, frozen-encoder, distribution-prior, and action-sensitive, and re-sense the stable-worldmodel benchmark so that only the observation differs from the image baselines. All three plan without collapse: the distribution-prior model is statistically equivalent to its re-evaluated image counterpart on every benchmark, and the action-sensitive model attains the strongest result in our controlled comparison where the most geometry moves. Probing explains why: object positions are almost perfectly linearly decodable and attention falls on the few moving points. Planning withstands heavy dropout never seen in training, though range noise defeats the thinnest scene. Geometry finally makes a commanded 3D target a natural goal interface: we construct the goal latent from the target and the current latent, at no cost in success rate, without a goal observation.
Problem

Research questions and friction points this paper is trying to address.

Latent Planning
Point Clouds
JEPA World Models
Geometric Observations
Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

point clouds
latent-space planning
JEPA world models
geometric observations
action-conditioned