Institution profile

Wayve

Industry researcheurope · gb
Official website
Research library12linked papers
Opportunities94open roles
Selected work

Representative Papers

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration

May 21, 2026

This work addresses the challenge of sparse rewards and long-horizon exploration in 3D environments, where existing curiosity-driven methods often fall into local loops or redundant exploration. The authors propose a novel approach that integrates an online 3D reconstruction-based world model with RGB sequence-based policy learning, introducing spatial persistence and episodic context into the curiosity mechanism for the first time. This enables the agent to distinguish genuinely novel regions and avoid revisiting uninformative areas. Relying solely on visual inputs, the method achieves zero-shot generalization from Habitat-Matterport 3D (HM3D) to both Gibson and AI-generated environments after pure curiosity-driven pre-training, significantly outperforming from-scratch baselines and demonstrating strong performance on downstream tasks such as apple picking and image-goal navigation.

0 citationsRead paper

LA-Pose: Latent Action Pretraining Meets Pose Estimation

Apr 30, 2026

This work addresses the heavy reliance of camera pose estimation on large-scale 3D-annotated data by proposing a self-supervised pretraining approach that, for the first time, integrates inverse and forward dynamics models into this task. By learning latent action representations from large-scale unlabeled driving videos and using them as input features to a fine-tuned pose estimator, the method substantially reduces dependence on annotated data. With only a small amount of high-quality 3D annotations, it achieves over a 10% improvement in pose accuracy compared to state-of-the-art feedforward methods on both the Waymo and PandaSet benchmarks, attaining leading performance with an order of magnitude fewer annotations.

0 citationsRead paper

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

Apr 28, 2026

Existing 3D scene reconstruction methods struggle to meet the demands of interactive graphics applications for high-quality, cross-view consistent semantic segmentation. This work proposes a novel approach based on Radiant Foam—a voxelized Voronoi grid—by introducing an explicit semantic feature field at the cell level and, for the first time, directly incorporating spatial regularization into this field to achieve joint spatial-semantic decomposition. This design significantly enhances cross-view consistency and effectively mitigates artifacts caused by occlusions and inconsistent supervision. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches such as Gaussian Grouping and SAGA in both object-level semantic segmentation accuracy and consistency.

0 citationsRead paper

CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering

Jul 11, 2025

Neural rendering struggles to simultaneously achieve efficiency, adaptivity, and representational fidelity. To address this, we propose Compressed Light Field Tokens (CLiFT): a compact, variable-length light field token representation enabling multi-fidelity rendering within a single network. Our method integrates a multi-view encoder, latent-space K-means clustering, a token compressor, and a pose-aware adaptive renderer that dynamically adjusts the number of tokens to balance computational cost and reconstruction quality. The key innovation is the first introduction of a scalable light field token mechanism, supporting fine-grained control over model complexity. Evaluated on RealEstate10K and DL3DV, CLiFT achieves state-of-the-art rendering quality while significantly reducing storage and computation overhead—delivering superior overall performance in terms of speed, quality, and memory efficiency.

0 citationsRead paper
Recent publications

Latest Papers

Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration

May 21, 2026

This work addresses the challenge of sparse rewards and long-horizon exploration in 3D environments, where existing curiosity-driven methods often fall into local loops or redundant exploration. The authors propose a novel approach that integrates an online 3D reconstruction-based world model with RGB sequence-based policy learning, introducing spatial persistence and episodic context into the curiosity mechanism for the first time. This enables the agent to distinguish genuinely novel regions and avoid revisiting uninformative areas. Relying solely on visual inputs, the method achieves zero-shot generalization from Habitat-Matterport 3D (HM3D) to both Gibson and AI-generated environments after pure curiosity-driven pre-training, significantly outperforming from-scratch baselines and demonstrating strong performance on downstream tasks such as apple picking and image-goal navigation.

0 citationsRead paper

LA-Pose: Latent Action Pretraining Meets Pose Estimation

Apr 30, 2026

This work addresses the heavy reliance of camera pose estimation on large-scale 3D-annotated data by proposing a self-supervised pretraining approach that, for the first time, integrates inverse and forward dynamics models into this task. By learning latent action representations from large-scale unlabeled driving videos and using them as input features to a fine-tuned pose estimator, the method substantially reduces dependence on annotated data. With only a small amount of high-quality 3D annotations, it achieves over a 10% improvement in pose accuracy compared to state-of-the-art feedforward methods on both the Waymo and PandaSet benchmarks, attaining leading performance with an order of magnitude fewer annotations.

0 citationsRead paper

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

Apr 28, 2026

Existing 3D scene reconstruction methods struggle to meet the demands of interactive graphics applications for high-quality, cross-view consistent semantic segmentation. This work proposes a novel approach based on Radiant Foam—a voxelized Voronoi grid—by introducing an explicit semantic feature field at the cell level and, for the first time, directly incorporating spatial regularization into this field to achieve joint spatial-semantic decomposition. This design significantly enhances cross-view consistency and effectively mitigates artifacts caused by occlusions and inconsistent supervision. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches such as Gaussian Grouping and SAGA in both object-level semantic segmentation accuracy and consistency.

0 citationsRead paper

CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering

Jul 11, 2025

Neural rendering struggles to simultaneously achieve efficiency, adaptivity, and representational fidelity. To address this, we propose Compressed Light Field Tokens (CLiFT): a compact, variable-length light field token representation enabling multi-fidelity rendering within a single network. Our method integrates a multi-view encoder, latent-space K-means clustering, a token compressor, and a pose-aware adaptive renderer that dynamically adjusts the number of tokens to balance computational cost and reconstruction quality. The key innovation is the first introduction of a scalable light field token mechanism, supporting fine-grained control over model complexity. Evaluated on RealEstate10K and DL3DV, CLiFT achieves state-of-the-art rendering quality while significantly reducing storage and computation overhead—delivering superior overall performance in terms of speed, quality, and memory efficiency.

0 citationsRead paper