Institution profile

Fraunhofer Institute for Vehicle Dynamics and Transport Systems IVI

Academic institutioneurope · de
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Road-Aware Anomaly Segmentation with Query-Guided Polygons and CLIP in Autonomous Driving

Jul 05, 2026

This work addresses the challenge of detecting road anomalies—such as unknown obstacles—in open-world autonomous driving, where semantic segmentation models often fail to recognize out-of-distribution objects. The authors propose a lightweight, post-processing framework that requires neither model retraining nor access to anomalous data. Their approach innovatively integrates query-guided polygonal road geometry priors with zero-shot CLIP-based semantic filtering, supporting both in-distribution and generalized out-of-distribution prompts. By leveraging confidence analysis from a Mask Transformer, the method enables unsupervised and interpretable precise segmentation of anomalous regions. Evaluated on Fishyscapes, SMIYC, and RoadAnomaly benchmarks, the framework significantly outperforms existing methods, achieving state-of-the-art average precision on the Fishyscapes LostAndFound subset and demonstrating strong deployment feasibility and robustness.

0 citationsRead paper

FlexPath: Learned Semantic Path Priors for Image-Based Planning

Jun 08, 2026

This work addresses the limited adaptability of existing learning-based path planning methods, which are typically constrained to shortest-path objectives and struggle to accommodate diverse task criteria. To overcome this, the authors propose FlexPath, a two-stage framework: the first stage leverages imitation learning to extract task-agnostic feasible path priors from visual maps, while the second introduces a differentiable Path Shape Objective (PSO) that efficiently adapts paths to varying task preferences without altering their underlying structure. FlexPath is the first approach to decouple path feasibility from task-specific objectives, enabling a single pre-trained model to generalize across multiple planning criteria through objective-level adjustments while remaining compatible with classical planners. Experiments show that on the TMP dataset, FlexPath reduces search overhead by 14.3% compared to TransPath, yields lower average path costs, successfully generalizes to three unseen domains in a zero-shot setting, and achieves a 96.8% success rate in full obstacle avoidance.

0 citationsRead paper

SOCC-ICP: Semantics-Assisted Odometry based on Occupancy Grids and ICP

May 14, 2026

This work addresses the performance degradation of LiDAR odometry and downstream tasks in unknown environments caused by inconsistent map representations. To this end, we propose a unified odometry framework that integrates semantic occupancy grid mapping with scan registration. Our method leverages voxel-based encoding of both geometric and semantic statistics to adaptively select between point-to-point and point-to-plane ICP strategies, while employing ray casting to remove dynamic objects. By unifying semantic occupancy grids and ICP-based odometry within a single representation for the first time, our approach eliminates redundant structures and maintains robustness even without semantic labels, achieving further accuracy gains when semantic information is available. Experimental results demonstrate state-of-the-art performance in both standard and geometrically degenerate scenarios.

0 citationsRead paper

AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards

May 20, 2025

Existing apple orchard monitoring datasets suffer from insufficient scene diversity, labor-intensive annotation, inadequate coverage of phenological growth stages, and lack of stereo imagery—limiting progress in fruit localization, yield estimation, and 3D reconstruction. To address these gaps, we introduce the first large-scale, binocular stereo image dataset spanning the complete apple growth cycle, systematically aligned with the BBCH phenological scale and enriched with dense pixel-level annotations and agronomically validated labels. We establish a standardized benchmark integrating agricultural science and computer vision, bridging critical gaps in growth modeling and 3D perception. Evaluation on this dataset demonstrates substantial improvements: YOLOv8 and Faster R-CNN achieve F1-score gains of 7.69% and 31.06%, respectively, in fruit detection; six-stage phenological classification accuracy exceeds 95%; and high-precision fruit localization and orchard-scale 3D reconstruction are enabled.

0 citationsRead paper
Recent publications

Latest Papers

Road-Aware Anomaly Segmentation with Query-Guided Polygons and CLIP in Autonomous Driving

Jul 05, 2026

This work addresses the challenge of detecting road anomalies—such as unknown obstacles—in open-world autonomous driving, where semantic segmentation models often fail to recognize out-of-distribution objects. The authors propose a lightweight, post-processing framework that requires neither model retraining nor access to anomalous data. Their approach innovatively integrates query-guided polygonal road geometry priors with zero-shot CLIP-based semantic filtering, supporting both in-distribution and generalized out-of-distribution prompts. By leveraging confidence analysis from a Mask Transformer, the method enables unsupervised and interpretable precise segmentation of anomalous regions. Evaluated on Fishyscapes, SMIYC, and RoadAnomaly benchmarks, the framework significantly outperforms existing methods, achieving state-of-the-art average precision on the Fishyscapes LostAndFound subset and demonstrating strong deployment feasibility and robustness.

0 citationsRead paper

FlexPath: Learned Semantic Path Priors for Image-Based Planning

Jun 08, 2026

This work addresses the limited adaptability of existing learning-based path planning methods, which are typically constrained to shortest-path objectives and struggle to accommodate diverse task criteria. To overcome this, the authors propose FlexPath, a two-stage framework: the first stage leverages imitation learning to extract task-agnostic feasible path priors from visual maps, while the second introduces a differentiable Path Shape Objective (PSO) that efficiently adapts paths to varying task preferences without altering their underlying structure. FlexPath is the first approach to decouple path feasibility from task-specific objectives, enabling a single pre-trained model to generalize across multiple planning criteria through objective-level adjustments while remaining compatible with classical planners. Experiments show that on the TMP dataset, FlexPath reduces search overhead by 14.3% compared to TransPath, yields lower average path costs, successfully generalizes to three unseen domains in a zero-shot setting, and achieves a 96.8% success rate in full obstacle avoidance.

0 citationsRead paper

SOCC-ICP: Semantics-Assisted Odometry based on Occupancy Grids and ICP

May 14, 2026

This work addresses the performance degradation of LiDAR odometry and downstream tasks in unknown environments caused by inconsistent map representations. To this end, we propose a unified odometry framework that integrates semantic occupancy grid mapping with scan registration. Our method leverages voxel-based encoding of both geometric and semantic statistics to adaptively select between point-to-point and point-to-plane ICP strategies, while employing ray casting to remove dynamic objects. By unifying semantic occupancy grids and ICP-based odometry within a single representation for the first time, our approach eliminates redundant structures and maintains robustness even without semantic labels, achieving further accuracy gains when semantic information is available. Experimental results demonstrate state-of-the-art performance in both standard and geometrically degenerate scenarios.

0 citationsRead paper

AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards

May 20, 2025

Existing apple orchard monitoring datasets suffer from insufficient scene diversity, labor-intensive annotation, inadequate coverage of phenological growth stages, and lack of stereo imagery—limiting progress in fruit localization, yield estimation, and 3D reconstruction. To address these gaps, we introduce the first large-scale, binocular stereo image dataset spanning the complete apple growth cycle, systematically aligned with the BBCH phenological scale and enriched with dense pixel-level annotations and agronomically validated labels. We establish a standardized benchmark integrating agricultural science and computer vision, bridging critical gaps in growth modeling and 3D perception. Evaluation on this dataset demonstrates substantial improvements: YOLOv8 and Faster R-CNN achieve F1-score gains of 7.69% and 31.06%, respectively, in fruit detection; six-stage phenological classification accuracy exceeds 95%; and high-precision fruit localization and orchard-scale 3D reconstruction are enabled.

0 citationsRead paper