Institution profile

SenseAuto Research

Industry researchasia · cn
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

Feb 03, 2026

This work addresses the challenge that existing driving world models struggle to simultaneously achieve instance-level temporal consistency and spatial geometric fidelity in multi-view video generation. To this end, the authors propose an instance-aware generative framework comprising an Instance Flow Guider to preserve cross-frame instance identity consistency and a Spatial Geometric Aligner to model spatial layout and occlusion relationships. By propagating instance features across frames, enforcing geometric alignment, and leveraging procedurally generated rare hazardous scenarios from the CARLA platform, the method achieves state-of-the-art video generation quality on the nuScenes dataset and significantly enhances the evaluation capability of autonomous driving systems in safety-critical scenarios.

6 citationsRead paper

Synthetic-to-Real Translation for Class-Agnostic Motion Prediction

Jul 07, 2026

This work addresses the challenges of high annotation costs for motion labels in real-world scenarios and performance degradation in motion prediction due to domain shift between synthetic and real data. To this end, the authors propose a dual-module transfer framework that integrates object-aware and object-assisted components, combining objectness-aware prior modeling, domain-invariant feature learning, and motion label denoising. Leveraging physics-based 4D LiDAR synthesis, they construct Motion4D—the first synthetic dataset specifically designed for motion prediction. Experimental results demonstrate that the proposed approach substantially narrows the domain gap between synthetic and real data, achieving robust and superior motion prediction performance in real-world settings.

0 citationsRead paper

DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

Feb 09, 2026

This work proposes DriveMamba, a novel end-to-end autonomous driving framework that addresses the limitations of existing systems—namely, information loss and error accumulation from modular designs, as well as the computational inefficiency of high-complexity attention mechanisms in modeling dynamic multi-task and multi-sensor relationships. DriveMamba introduces, for the first time, a linear-complexity state space model into end-to-end driving, featuring a unified single-stage Mamba decoder. It leverages task-centric sparse token representations, 3D spatial position ordering, and a bidirectional trajectory-guided “local-to-global” scanning strategy to enable dynamic task relationship modeling, implicit view alignment, and long-range temporal fusion. Experiments on nuScenes and Bench2Drive demonstrate that DriveMamba significantly outperforms current methods in terms of performance, generalization, and computational efficiency.

0 citationsRead paper

FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers

Oct 17, 2025

In multi-view 2D image-based 3D object detection, inaccurate depth estimation—manifesting as depth discontinuities at object boundaries, poor small-object discrimination, cross-view inconsistency, and scale sensitivity—remains a critical challenge. To address these issues, this paper proposes the Frequency-Aware Positional Depth Embedding (FAPDE) framework. Its key contributions are: (1) a Frequency-Aware Spatial Pyramid Encoder (FSPE) that jointly encodes multi-level high-frequency edge and low-frequency semantic features to enhance structural fidelity in depth prediction; (2) a Cross-View Scale-Invariant Depth Predictor (CSDP) that jointly optimizes depth consistency across views and robustness to object scale variations; and (3) a hybrid depth supervision scheme coupled with channel-attention-guided 2D–3D feature fusion. Evaluated on nuScenes, FAPDE achieves state-of-the-art performance, notably improving depth continuity and small-object recall while boosting overall 3D detection accuracy.

0 citationsRead paper

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs

Oct 01, 2025

Search-augmented large language models (LLMs) suffer from insufficient robustness in multi-hop reasoning due to decomposition errors, retrieval failures, and inference mistakes. To address this, we propose EraseRL, a novel erasable reinforcement learning framework. EraseRL introduces, for the first time, a dynamic “identify–erase–locally regenerate” mechanism within the reasoning chain, enabling real-time localization and correction of erroneous reasoning steps to halt error propagation—thereby shifting the paradigm from fragile to resilient reasoning. The method integrates search-augmented architecture with end-to-end reinforcement learning, optimizing full-chain reliability without requiring additional human annotations. Evaluated on HotpotQA and MuSiQue benchmarks, EraseRL improves exact match (EM) by +8.48% and +5.38%, and F1 by +11.56% and +7.22%, respectively, for 3B- and 7B-scale models—significantly surpassing state-of-the-art approaches.

0 citationsRead paper
Recent publications

Latest Papers

Synthetic-to-Real Translation for Class-Agnostic Motion Prediction

Jul 07, 2026

This work addresses the challenges of high annotation costs for motion labels in real-world scenarios and performance degradation in motion prediction due to domain shift between synthetic and real data. To this end, the authors propose a dual-module transfer framework that integrates object-aware and object-assisted components, combining objectness-aware prior modeling, domain-invariant feature learning, and motion label denoising. Leveraging physics-based 4D LiDAR synthesis, they construct Motion4D—the first synthetic dataset specifically designed for motion prediction. Experimental results demonstrate that the proposed approach substantially narrows the domain gap between synthetic and real data, achieving robust and superior motion prediction performance in real-world settings.

0 citationsRead paper

DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

Feb 09, 2026

This work proposes DriveMamba, a novel end-to-end autonomous driving framework that addresses the limitations of existing systems—namely, information loss and error accumulation from modular designs, as well as the computational inefficiency of high-complexity attention mechanisms in modeling dynamic multi-task and multi-sensor relationships. DriveMamba introduces, for the first time, a linear-complexity state space model into end-to-end driving, featuring a unified single-stage Mamba decoder. It leverages task-centric sparse token representations, 3D spatial position ordering, and a bidirectional trajectory-guided “local-to-global” scanning strategy to enable dynamic task relationship modeling, implicit view alignment, and long-range temporal fusion. Experiments on nuScenes and Bench2Drive demonstrate that DriveMamba significantly outperforms current methods in terms of performance, generalization, and computational efficiency.

0 citationsRead paper

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

Feb 03, 2026

This work addresses the challenge that existing driving world models struggle to simultaneously achieve instance-level temporal consistency and spatial geometric fidelity in multi-view video generation. To this end, the authors propose an instance-aware generative framework comprising an Instance Flow Guider to preserve cross-frame instance identity consistency and a Spatial Geometric Aligner to model spatial layout and occlusion relationships. By propagating instance features across frames, enforcing geometric alignment, and leveraging procedurally generated rare hazardous scenarios from the CARLA platform, the method achieves state-of-the-art video generation quality on the nuScenes dataset and significantly enhances the evaluation capability of autonomous driving systems in safety-critical scenarios.

6 citationsRead paper

FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers

Oct 17, 2025

In multi-view 2D image-based 3D object detection, inaccurate depth estimation—manifesting as depth discontinuities at object boundaries, poor small-object discrimination, cross-view inconsistency, and scale sensitivity—remains a critical challenge. To address these issues, this paper proposes the Frequency-Aware Positional Depth Embedding (FAPDE) framework. Its key contributions are: (1) a Frequency-Aware Spatial Pyramid Encoder (FSPE) that jointly encodes multi-level high-frequency edge and low-frequency semantic features to enhance structural fidelity in depth prediction; (2) a Cross-View Scale-Invariant Depth Predictor (CSDP) that jointly optimizes depth consistency across views and robustness to object scale variations; and (3) a hybrid depth supervision scheme coupled with channel-attention-guided 2D–3D feature fusion. Evaluated on nuScenes, FAPDE achieves state-of-the-art performance, notably improving depth continuity and small-object recall while boosting overall 3D detection accuracy.

0 citationsRead paper

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs

Oct 01, 2025

Search-augmented large language models (LLMs) suffer from insufficient robustness in multi-hop reasoning due to decomposition errors, retrieval failures, and inference mistakes. To address this, we propose EraseRL, a novel erasable reinforcement learning framework. EraseRL introduces, for the first time, a dynamic “identify–erase–locally regenerate” mechanism within the reasoning chain, enabling real-time localization and correction of erroneous reasoning steps to halt error propagation—thereby shifting the paradigm from fragile to resilient reasoning. The method integrates search-augmented architecture with end-to-end reinforcement learning, optimizing full-chain reliability without requiring additional human annotations. Evaluated on HotpotQA and MuSiQue benchmarks, EraseRL improves exact match (EM) by +8.48% and +5.38%, and F1 by +11.56% and +7.22%, respectively, for 3B- and 7B-scale models—significantly surpassing state-of-the-art approaches.

0 citationsRead paper