DriftParking: Trajectory Modeling via Drifting Field for End-to-End Automated Parking
为解决自动泊车中轨迹生成的精度和效率问题,提出DriftParking框架,通过漂移场模型和结构化偏差监督生成高质量轨迹。
为解决自动泊车中轨迹生成的精度和效率问题,提出DriftParking框架,通过漂移场模型和结构化偏差监督生成高质量轨迹。
本文针对遮挡边界估计的监督不足问题,通过创建RealOOB基准,提供一致定义的几何基础标签,评估多种边界和深度估计器性能。
This work addresses the significant performance degradation of semantic segmentation models trained on synthetic near-infrared (NIR) images when deployed in real-world automotive scenarios, primarily caused by domain shift and scarce annotations. To mitigate this, the authors propose a generative augmentation framework comprising Target Style Adaptation (TSA) and Voronoi Style Diversification (VSD). TSA leverages low-rank fine-tuned diffusion models with structure-preserving multi-signal control to translate synthetic images into realistic NIR styles, while VSD employs geometry-guided disentanglement to decouple texture and shape inductive biases. As the first systematic study on synthetic-to-real domain adaptation for automotive NIR imagery, the proposed method reduces domain gaps by 63.6% and 28.4% in internal and external scenes, respectively, substantially enhancing the robustness and generalization of diverse segmentation models in real-world conditions.
This work addresses the significant performance degradation of conventional RGB-based object detection under extremely low-light conditions. To overcome this limitation, the authors propose a dual-stream fusion framework that integrates CLAHE-enhanced RGB images with voxelized event data. Central to their approach is an adaptive cross-modal attention mechanism grounded in minimum-variance linear estimation theory, which asymptotically approximates the Gauss–Markov optimal fusion weights. The study further establishes, for the first time, theoretical bounds relating the conservation properties of event voxelization to its temporal resolution. Evaluated on the LLE-VOS benchmark, the method achieves 65.54% recall, 53.85% precision, and 59.12% F1-score, substantially outperforming single-modality approaches and demonstrating both its illumination-adaptive capability and theoretical soundness.
This work addresses the challenge of inaccurate relative camera pose estimation in in-cabin fisheye imaging, which suffers from severe distortion and spatial constraints. The authors propose a single-pass Transformer-based architecture that leverages a frozen DINOv2 backbone for feature extraction and a ViT-Small decoder to model geometric correspondences between reference and target images, directly regressing metric-scale translation and rotation. Notably, the model is trained exclusively on synthetic data yet achieves strong domain generalization to real in-cabin scenes without requiring known camera intrinsics, yielding physically plausible poses. Evaluated on both the newly introduced In-Cabin-Pose benchmark and the 7-Scenes dataset, the method demonstrates high accuracy and real-time performance, with code and dataset publicly released to support safety-critical driver monitoring applications.
为解决自动泊车中轨迹生成的精度和效率问题,提出DriftParking框架,通过漂移场模型和结构化偏差监督生成高质量轨迹。
本文针对遮挡边界估计的监督不足问题,通过创建RealOOB基准,提供一致定义的几何基础标签,评估多种边界和深度估计器性能。
This work addresses the significant performance degradation of semantic segmentation models trained on synthetic near-infrared (NIR) images when deployed in real-world automotive scenarios, primarily caused by domain shift and scarce annotations. To mitigate this, the authors propose a generative augmentation framework comprising Target Style Adaptation (TSA) and Voronoi Style Diversification (VSD). TSA leverages low-rank fine-tuned diffusion models with structure-preserving multi-signal control to translate synthetic images into realistic NIR styles, while VSD employs geometry-guided disentanglement to decouple texture and shape inductive biases. As the first systematic study on synthetic-to-real domain adaptation for automotive NIR imagery, the proposed method reduces domain gaps by 63.6% and 28.4% in internal and external scenes, respectively, substantially enhancing the robustness and generalization of diverse segmentation models in real-world conditions.
This work addresses the significant performance degradation of conventional RGB-based object detection under extremely low-light conditions. To overcome this limitation, the authors propose a dual-stream fusion framework that integrates CLAHE-enhanced RGB images with voxelized event data. Central to their approach is an adaptive cross-modal attention mechanism grounded in minimum-variance linear estimation theory, which asymptotically approximates the Gauss–Markov optimal fusion weights. The study further establishes, for the first time, theoretical bounds relating the conservation properties of event voxelization to its temporal resolution. Evaluated on the LLE-VOS benchmark, the method achieves 65.54% recall, 53.85% precision, and 59.12% F1-score, substantially outperforming single-modality approaches and demonstrating both its illumination-adaptive capability and theoretical soundness.
This work addresses the challenge of inaccurate relative camera pose estimation in in-cabin fisheye imaging, which suffers from severe distortion and spatial constraints. The authors propose a single-pass Transformer-based architecture that leverages a frozen DINOv2 backbone for feature extraction and a ViT-Small decoder to model geometric correspondences between reference and target images, directly regressing metric-scale translation and rotation. Notably, the model is trained exclusively on synthetic data yet achieves strong domain generalization to real in-cabin scenes without requiring known camera intrinsics, yielding physically plausible poses. Evaluated on both the newly introduced In-Cabin-Pose benchmark and the 7-Scenes dataset, the method demonstrates high accuracy and real-time performance, with code and dataset publicly released to support safety-critical driver monitoring applications.