Institution profile

aiMotive Kft

Industry researcheurope · hu
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Multimodal Scenario Similarity Search for Autonomous Driving

Jul 10, 2026

This work addresses the absence of efficient multimodal scene similarity retrieval methods in large-scale autonomous driving datasets that jointly account for visual appearance and dynamic behavior. The authors propose a unified multimodal retrieval framework that, for the first time, systematically integrates explicit trajectory matching (Exo-Trajectory) with Transformer-based trajectory representations (ScenarioFormer), alongside vision embeddings trained via contrastive learning. Experimental results demonstrate that trajectory-based representations excel in dynamic scenarios such as cut-ins and turns, while visual methods are better suited for appearance-dominated scenes. Crucially, multimodal fusion substantially enhances overall retrieval performance, revealing the complementary nature of appearance and motion features in assessing scene similarity.

0 citationsRead paper

CCLSTM: Coupled Convolutional Long-Short Term Memory Network for Occupancy Flow Forecasting

Jun 06, 2025

Accurate future state prediction of dynamic agents in autonomous driving heavily relies on high-quality vectorized inputs and computationally intensive Transformer architectures. Method: This paper proposes an end-to-end lightweight motion modeling approach based on Occupancy Flow Fields (OFF), eliminating both vectorization preprocessing and self-attention mechanisms. It introduces a novel, purely convolutional coupled LSTM architecture that jointly models spatiotemporal dependencies via trainable convolutional recurrence. The method adopts a direct occupancy flow field encoding–decoding paradigm, drastically reducing computational overhead and deployment cost. Contribution/Results: Evaluated on the 2024 Waymo Occupancy Flow Prediction Challenge, our method achieves state-of-the-art (SOTA) performance, ranking first across all official metrics.

0 citationsRead paper

Hybrid Rendering for Multimodal Autonomous Driving: Merging Neural and Physics-Based Simulation

Mar 12, 2025

To address the limited generalization and poor physical controllability of neural reconstruction in autonomous driving simulation, this paper proposes a hybrid framework integrating neural and physics-based rendering. Methodologically, we introduce NeRF2GS—a novel distillation paradigm where a depth-regularized Neural Radiance Field (NeRF) serves as the teacher model and 3D Gaussian Splatting (3DGS) as the student. We incorporate LiDAR point cloud–robust depth supervision, block-parallel optimization, and depth-aware compositing to enable reconstruction of scenes spanning hundreds of square kilometers, while jointly outputting RGB, semantic segmentation, surface normals, and depth maps. Our contributions include the first demonstration of arbitrary placement of dynamic agents, adjustable environmental parameters, and multi-view real-time rendering (>30 FPS). This significantly improves novel-view synthesis quality for road surfaces and lane markings, and ensures compatibility with multi-camera setups and high-fidelity LiDAR simulation.

0 citationsRead paper
Recent publications

Latest Papers

Multimodal Scenario Similarity Search for Autonomous Driving

Jul 10, 2026

This work addresses the absence of efficient multimodal scene similarity retrieval methods in large-scale autonomous driving datasets that jointly account for visual appearance and dynamic behavior. The authors propose a unified multimodal retrieval framework that, for the first time, systematically integrates explicit trajectory matching (Exo-Trajectory) with Transformer-based trajectory representations (ScenarioFormer), alongside vision embeddings trained via contrastive learning. Experimental results demonstrate that trajectory-based representations excel in dynamic scenarios such as cut-ins and turns, while visual methods are better suited for appearance-dominated scenes. Crucially, multimodal fusion substantially enhances overall retrieval performance, revealing the complementary nature of appearance and motion features in assessing scene similarity.

0 citationsRead paper

CCLSTM: Coupled Convolutional Long-Short Term Memory Network for Occupancy Flow Forecasting

Jun 06, 2025

Accurate future state prediction of dynamic agents in autonomous driving heavily relies on high-quality vectorized inputs and computationally intensive Transformer architectures. Method: This paper proposes an end-to-end lightweight motion modeling approach based on Occupancy Flow Fields (OFF), eliminating both vectorization preprocessing and self-attention mechanisms. It introduces a novel, purely convolutional coupled LSTM architecture that jointly models spatiotemporal dependencies via trainable convolutional recurrence. The method adopts a direct occupancy flow field encoding–decoding paradigm, drastically reducing computational overhead and deployment cost. Contribution/Results: Evaluated on the 2024 Waymo Occupancy Flow Prediction Challenge, our method achieves state-of-the-art (SOTA) performance, ranking first across all official metrics.

0 citationsRead paper

Hybrid Rendering for Multimodal Autonomous Driving: Merging Neural and Physics-Based Simulation

Mar 12, 2025

To address the limited generalization and poor physical controllability of neural reconstruction in autonomous driving simulation, this paper proposes a hybrid framework integrating neural and physics-based rendering. Methodologically, we introduce NeRF2GS—a novel distillation paradigm where a depth-regularized Neural Radiance Field (NeRF) serves as the teacher model and 3D Gaussian Splatting (3DGS) as the student. We incorporate LiDAR point cloud–robust depth supervision, block-parallel optimization, and depth-aware compositing to enable reconstruction of scenes spanning hundreds of square kilometers, while jointly outputting RGB, semantic segmentation, surface normals, and depth maps. Our contributions include the first demonstration of arbitrary placement of dynamic agents, adjustable environmental parameters, and multi-view real-time rendering (>30 FPS). This significantly improves novel-view synthesis quality for road surfaces and lane markings, and ensures compatibility with multi-camera setups and high-fidelity LiDAR simulation.

0 citationsRead paper