Accuracy- and Real-Time-Aware 4D Radar Preprocessing for Autonomous Driving Perception Systems
为提高4D雷达在自动驾驶感知系统中的应用,提出了一种包括P3DP、MF-KDE和ENS的预处理框架,以兼顾准确性、实时性和计算复杂度。
为提高4D雷达在自动驾驶感知系统中的应用,提出了一种包括P3DP、MF-KDE和ENS的预处理框架,以兼顾准确性、实时性和计算复杂度。
This work addresses the significant performance gap between Vision Transformers (ViTs) and convolutional neural networks (CNNs) in data-scarce scenarios, where existing knowledge distillation methods struggle to effectively transfer the inductive bias of CNNs. To this end, the authors propose iBKD, a novel framework that, for the first time, preserves and leverages the spatial grid structure of the teacher CNN throughout the entire ViT distillation process. By aligning grid-structured features between student and teacher via an inductive bias attention module—augmented with channel attention, deformable spatial attention, and a convolutional cross-attention mechanism used only during training—the method efficiently injects locality priors without introducing inference overhead. Evaluated across seven ViT backbones and six data-scarce benchmarks, iBKD consistently outperforms both locality-aware and general-purpose distillation approaches, with performance gains becoming more pronounced as data availability decreases.
Existing robotic policy learning approaches struggle to accommodate the diversity of user preferences: single-policy methods overlook preference heterogeneity, while user-specific alignment suffers from sparse and noisy feedback as well as high validation costs. To address these limitations, this work proposes the PREC framework, which jointly optimizes user preference clustering and reward learning. PREC first employs a shared trajectory encoder to extract universal representations, then simultaneously performs user clustering and cluster-wise reward modeling using preference labels. This enables the optimization of a dedicated policy for each cluster, preserving preference diversity while mitigating label sparsity and noise, and substantially reducing pre-deployment validation burden. Experiments in simulated locomotion environments demonstrate that PREC more accurately identifies groups of users with consistent preferences and yields policies that outperform single-policy baselines—and even surpass user-specific alignment—across three social welfare metrics.
This work addresses the challenge of maintaining consistent long-term LiDAR maps by unifying dynamic object removal and change detection—tasks typically treated independently in existing approaches. To this end, we propose MTD-Map, a single-stage map maintenance framework that jointly handles both tasks through explicit modeling of a mixed transition distribution (MTD), eliminating the need for task-specific modules. The method recursively encodes historical occupancy states to capture high-order temporal dependencies and incorporates a stability-driven adaptive update mechanism that effectively suppresses noise while preserving quasi-static structures. Experimental results demonstrate that MTD-Map achieves state-of-the-art performance in both dynamic object removal and change detection, while significantly reducing computational overhead, thereby validating its robustness and efficiency.
Existing multimodal trajectory prediction methods often neglect lane topology constraints, leading to the generation of infeasible trajectories under low-probability modes and thereby compromising autonomous driving safety. To address this issue, this work proposes the LAMP framework, which innovatively integrates lane topology into motion primitive learning. Specifically, a VQ-VAE is employed to construct shape-aware discrete intention representations, and a feasibility-aware intention selection mechanism is designed by incorporating lane priors. Trajectories are then generated via an attention-based decoder. Evaluated on Argoverse 2, the proposed method achieves state-of-the-art prediction accuracy while significantly improving both the physical feasibility and diversity of predicted trajectories.
为提高4D雷达在自动驾驶感知系统中的应用,提出了一种包括P3DP、MF-KDE和ENS的预处理框架,以兼顾准确性、实时性和计算复杂度。
This work addresses the significant performance gap between Vision Transformers (ViTs) and convolutional neural networks (CNNs) in data-scarce scenarios, where existing knowledge distillation methods struggle to effectively transfer the inductive bias of CNNs. To this end, the authors propose iBKD, a novel framework that, for the first time, preserves and leverages the spatial grid structure of the teacher CNN throughout the entire ViT distillation process. By aligning grid-structured features between student and teacher via an inductive bias attention module—augmented with channel attention, deformable spatial attention, and a convolutional cross-attention mechanism used only during training—the method efficiently injects locality priors without introducing inference overhead. Evaluated across seven ViT backbones and six data-scarce benchmarks, iBKD consistently outperforms both locality-aware and general-purpose distillation approaches, with performance gains becoming more pronounced as data availability decreases.
Existing robotic policy learning approaches struggle to accommodate the diversity of user preferences: single-policy methods overlook preference heterogeneity, while user-specific alignment suffers from sparse and noisy feedback as well as high validation costs. To address these limitations, this work proposes the PREC framework, which jointly optimizes user preference clustering and reward learning. PREC first employs a shared trajectory encoder to extract universal representations, then simultaneously performs user clustering and cluster-wise reward modeling using preference labels. This enables the optimization of a dedicated policy for each cluster, preserving preference diversity while mitigating label sparsity and noise, and substantially reducing pre-deployment validation burden. Experiments in simulated locomotion environments demonstrate that PREC more accurately identifies groups of users with consistent preferences and yields policies that outperform single-policy baselines—and even surpass user-specific alignment—across three social welfare metrics.
This work addresses the challenge of maintaining consistent long-term LiDAR maps by unifying dynamic object removal and change detection—tasks typically treated independently in existing approaches. To this end, we propose MTD-Map, a single-stage map maintenance framework that jointly handles both tasks through explicit modeling of a mixed transition distribution (MTD), eliminating the need for task-specific modules. The method recursively encodes historical occupancy states to capture high-order temporal dependencies and incorporates a stability-driven adaptive update mechanism that effectively suppresses noise while preserving quasi-static structures. Experimental results demonstrate that MTD-Map achieves state-of-the-art performance in both dynamic object removal and change detection, while significantly reducing computational overhead, thereby validating its robustness and efficiency.
Existing multimodal trajectory prediction methods often neglect lane topology constraints, leading to the generation of infeasible trajectories under low-probability modes and thereby compromising autonomous driving safety. To address this issue, this work proposes the LAMP framework, which innovatively integrates lane topology into motion primitive learning. Specifically, a VQ-VAE is employed to construct shape-aware discrete intention representations, and a feasibility-aware intention selection mechanism is designed by incorporating lane priors. Trajectories are then generated via an attention-based decoder. Evaluated on Argoverse 2, the proposed method achieves state-of-the-art prediction accuracy while significantly improving both the physical feasibility and diversity of predicted trajectories.