Institution profile

Mannheim University of Applied Sciences

Academic institutioneurope · de
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

May 13, 2026

Real-world data often fail to meet the demand for high-quality, diverse pedestrian datasets required for autonomous driving, particularly in safety-critical scenarios. This work proposes a StyleGAN2-based controllable generation method that automatically maps facial textures and identity-level appearance variations onto a unified 3D pedestrian mesh, enabling efficient synthesis of large-scale, simulation-ready assets without requiring new geometric modeling. To the best of our knowledge, this is the first application of controllable generative AI to diversify 3D pedestrian appearances. Experiments demonstrate that the synthesized data significantly enhance the robustness of 2D detection models, while also revealing the high sensitivity of 3D perception models to geometric domain discrepancies—providing critical insights for simulation-based training.

0 citationsRead paper

TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

Apr 29, 2026

Existing methods for detecting AI-generated images suffer from limited generalization and fail to fully leverage the potential of modern vision foundation models (VFMs). This work presents the first systematic evaluation of various VFMs in their out-of-the-box performance on detecting both AI-generated and AI-edited images. To enhance feature aggregation, the study introduces a lightweight classification head incorporating tunable attention pooling (TAP). Experimental results demonstrate that the proposed approach significantly improves detection accuracy, surpassing the original CLIP model by over 12% across multiple benchmarks and achieving new state-of-the-art performance on two challenging in-the-wild detection datasets.

0 citationsRead paper

PointTransformerX:Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms

Apr 27, 2026

This work addresses the challenge of deploying 3D point cloud perception models on non-NVIDIA platforms by proposing the first fully native PyTorch implementation of a 3D point cloud Vision Transformer backbone, eliminating all custom CUDA operators and external dependencies. The architecture directly models spatial relationships through self-attention without explicit neighborhood construction, incorporates 3D-GS-RoPE for geometry-aware positional encoding, replaces sparse convolutions with linear projection for patch embedding, and introduces a lightweight feed-forward network alongside a dynamic attention window scaling mechanism. Evaluated on ScanNet, the model achieves 98.7% of PointTransformer V3’s accuracy while reducing parameters by 79.2%, accelerating inference by 1.6×, occupying only 253 MB of memory, and enabling seamless cross-platform execution on NVIDIA GPUs, AMD GPUs (via ROCm), and CPUs.

0 citationsRead paper

SSFT: A Lightweight Spectral-Spatial Fusion Transformer for Generic Hyperspectral Classification

Apr 17, 2026

This work addresses the challenges of hyperspectral image classification—particularly high dimensionality, spectral redundancy, scarce annotations, and domain shift—which are exacerbated in non-remote-sensing scenarios by data sparsity and class imbalance. To tackle these issues, the authors propose a lightweight Spectral-Spatial Fusion Transformer (SSFT), which, for the first time, efficiently integrates decoupled spectral and spatial modeling through a cross-attention mechanism. Remarkably, SSFT achieves robust training without data augmentation while using fewer than 2% of the parameters of previous state-of-the-art methods. The model attains top performance on the HSI-Benchmark and remains competitive on the larger-scale SpectralEarth benchmark, demonstrating an exceptional balance between computational efficiency and generalization capability.

0 citationsRead paper

D-PLS: Decoupled Semantic Segmentation for 4D-Panoptic-LiDAR-Segmentation

Jan 27, 2025

In 4D point cloud spatiotemporal panoptic segmentation, tightly coupled semantic classification and instance discrimination limit overall performance. Method: We propose a decoupled 4D panoptic LiDAR segmentation framework that explicitly separates semantic and instance tasks: single-frame semantic predictions serve as temporal guidance signals to drive a lightweight instance grouping module, without modifying or retraining existing semantic backbones. The pipeline—comprising single-scan semantic segmentation, temporal aggregation, and guidance-driven instance grouping—is modular and backbone-agnostic. Contribution/Results: Our plug-and-play design improves single-frame semantic accuracy while enabling end-to-end evaluation on SemanticKITTI using the LSTQ metric. Experiments demonstrate consistent and comprehensive improvements across all LSTQ sub-metrics over strong baselines, validating the synergistic gains of decoupling classification and association tasks.

0 citationsRead paper
Recent publications

Latest Papers

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

May 13, 2026

Real-world data often fail to meet the demand for high-quality, diverse pedestrian datasets required for autonomous driving, particularly in safety-critical scenarios. This work proposes a StyleGAN2-based controllable generation method that automatically maps facial textures and identity-level appearance variations onto a unified 3D pedestrian mesh, enabling efficient synthesis of large-scale, simulation-ready assets without requiring new geometric modeling. To the best of our knowledge, this is the first application of controllable generative AI to diversify 3D pedestrian appearances. Experiments demonstrate that the synthesized data significantly enhance the robustness of 2D detection models, while also revealing the high sensitivity of 3D perception models to geometric domain discrepancies—providing critical insights for simulation-based training.

0 citationsRead paper

TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

Apr 29, 2026

Existing methods for detecting AI-generated images suffer from limited generalization and fail to fully leverage the potential of modern vision foundation models (VFMs). This work presents the first systematic evaluation of various VFMs in their out-of-the-box performance on detecting both AI-generated and AI-edited images. To enhance feature aggregation, the study introduces a lightweight classification head incorporating tunable attention pooling (TAP). Experimental results demonstrate that the proposed approach significantly improves detection accuracy, surpassing the original CLIP model by over 12% across multiple benchmarks and achieving new state-of-the-art performance on two challenging in-the-wild detection datasets.

0 citationsRead paper

PointTransformerX:Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms

Apr 27, 2026

This work addresses the challenge of deploying 3D point cloud perception models on non-NVIDIA platforms by proposing the first fully native PyTorch implementation of a 3D point cloud Vision Transformer backbone, eliminating all custom CUDA operators and external dependencies. The architecture directly models spatial relationships through self-attention without explicit neighborhood construction, incorporates 3D-GS-RoPE for geometry-aware positional encoding, replaces sparse convolutions with linear projection for patch embedding, and introduces a lightweight feed-forward network alongside a dynamic attention window scaling mechanism. Evaluated on ScanNet, the model achieves 98.7% of PointTransformer V3’s accuracy while reducing parameters by 79.2%, accelerating inference by 1.6×, occupying only 253 MB of memory, and enabling seamless cross-platform execution on NVIDIA GPUs, AMD GPUs (via ROCm), and CPUs.

0 citationsRead paper

SSFT: A Lightweight Spectral-Spatial Fusion Transformer for Generic Hyperspectral Classification

Apr 17, 2026

This work addresses the challenges of hyperspectral image classification—particularly high dimensionality, spectral redundancy, scarce annotations, and domain shift—which are exacerbated in non-remote-sensing scenarios by data sparsity and class imbalance. To tackle these issues, the authors propose a lightweight Spectral-Spatial Fusion Transformer (SSFT), which, for the first time, efficiently integrates decoupled spectral and spatial modeling through a cross-attention mechanism. Remarkably, SSFT achieves robust training without data augmentation while using fewer than 2% of the parameters of previous state-of-the-art methods. The model attains top performance on the HSI-Benchmark and remains competitive on the larger-scale SpectralEarth benchmark, demonstrating an exceptional balance between computational efficiency and generalization capability.

0 citationsRead paper

D-PLS: Decoupled Semantic Segmentation for 4D-Panoptic-LiDAR-Segmentation

Jan 27, 2025

In 4D point cloud spatiotemporal panoptic segmentation, tightly coupled semantic classification and instance discrimination limit overall performance. Method: We propose a decoupled 4D panoptic LiDAR segmentation framework that explicitly separates semantic and instance tasks: single-frame semantic predictions serve as temporal guidance signals to drive a lightweight instance grouping module, without modifying or retraining existing semantic backbones. The pipeline—comprising single-scan semantic segmentation, temporal aggregation, and guidance-driven instance grouping—is modular and backbone-agnostic. Contribution/Results: Our plug-and-play design improves single-frame semantic accuracy while enabling end-to-end evaluation on SemanticKITTI using the LSTQ metric. Experiments demonstrate consistent and comprehensive improvements across all LSTQ sub-metrics over strong baselines, validating the synergistic gains of decoupling classification and association tasks.

0 citationsRead paper