Institution profile

Fraunhofer HHI

Academic institutioneurope · de
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs

Jun 18, 2026

This work addresses the challenge of simultaneously achieving substantial parameter reduction and performance preservation in large language model (LLM) compression by proposing an activation- and influence-aware low-rank approximation method. The approach uniquely integrates element-wise backward influence metrics into singular value decomposition (SVD)-based compression and employs a single closed-form alternating least squares (ALS) step to preserve functionality, offering both locality and monotonic descent properties. It is also orthogonal to end-to-end fine-tuning techniques. Experimental results demonstrate that with ≤60% of the original parameters retained, the method improves perplexity by over 18% compared to SVD-LLM(W). Moreover, it achieves comparable model quality using only ~10% calibration data while significantly reducing FLOPs, peak memory usage, and per-token latency.

0 citationsRead paper

Knowledge-Guided Failure Prediction: Detecting When Object Detectors Miss Safety-Critical Objects

Mar 26, 2026

This work addresses the critical challenge of silent failures in object detectors—such as missed pedestrian detections in safety-critical scenarios—that evade conventional out-of-distribution (OOD) detection methods. To this end, the authors propose KGFP, a knowledge-guided failure prediction framework that formulates detector failures as semantic inconsistencies between the detector’s internal features and embeddings from a vision foundation model. Leveraging a dual-encoder architecture and angular distance metrics, KGFP establishes a runtime selective prediction gating mechanism. Evaluated on pedestrian detection within the COCO benchmark, KGFP improves recall from 64.3% to 84.5% at a 5% false positive rate and consistently outperforms existing OOD detection approaches across six COCO-O domains.

0 citationsRead paper

Monocular 3D Object Position Estimation with VLMs for Human-Robot Interaction

Mar 01, 2026

This study addresses the challenge of high-precision 3D object localization for human-robot interaction by integrating monocular RGB images, natural language instructions, and robot state information. To this end, the authors propose an end-to-end framework built upon a pretrained vision-language model (VLM), enhanced with QLoRA-based efficient fine-tuning, a custom regression head, and a conditional routing mechanism. This design preserves the VLM’s general visual understanding capabilities while introducing dedicated 3D localization functionality. The work introduces a heterogeneous dataset comprising over 100,000 samples and demonstrates strong empirical performance, achieving a median absolute error of 13 mm—representing a fivefold improvement over the unmodified baseline. Notably, approximately 25% of predictions meet the accuracy threshold required for direct robotic manipulation.

0 citationsRead paper
Recent publications

Latest Papers

Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs

Jun 18, 2026

This work addresses the challenge of simultaneously achieving substantial parameter reduction and performance preservation in large language model (LLM) compression by proposing an activation- and influence-aware low-rank approximation method. The approach uniquely integrates element-wise backward influence metrics into singular value decomposition (SVD)-based compression and employs a single closed-form alternating least squares (ALS) step to preserve functionality, offering both locality and monotonic descent properties. It is also orthogonal to end-to-end fine-tuning techniques. Experimental results demonstrate that with ≤60% of the original parameters retained, the method improves perplexity by over 18% compared to SVD-LLM(W). Moreover, it achieves comparable model quality using only ~10% calibration data while significantly reducing FLOPs, peak memory usage, and per-token latency.

0 citationsRead paper

Knowledge-Guided Failure Prediction: Detecting When Object Detectors Miss Safety-Critical Objects

Mar 26, 2026

This work addresses the critical challenge of silent failures in object detectors—such as missed pedestrian detections in safety-critical scenarios—that evade conventional out-of-distribution (OOD) detection methods. To this end, the authors propose KGFP, a knowledge-guided failure prediction framework that formulates detector failures as semantic inconsistencies between the detector’s internal features and embeddings from a vision foundation model. Leveraging a dual-encoder architecture and angular distance metrics, KGFP establishes a runtime selective prediction gating mechanism. Evaluated on pedestrian detection within the COCO benchmark, KGFP improves recall from 64.3% to 84.5% at a 5% false positive rate and consistently outperforms existing OOD detection approaches across six COCO-O domains.

0 citationsRead paper

Monocular 3D Object Position Estimation with VLMs for Human-Robot Interaction

Mar 01, 2026

This study addresses the challenge of high-precision 3D object localization for human-robot interaction by integrating monocular RGB images, natural language instructions, and robot state information. To this end, the authors propose an end-to-end framework built upon a pretrained vision-language model (VLM), enhanced with QLoRA-based efficient fine-tuning, a custom regression head, and a conditional routing mechanism. This design preserves the VLM’s general visual understanding capabilities while introducing dedicated 3D localization functionality. The work introduces a heterogeneous dataset comprising over 100,000 samples and demonstrates strong empirical performance, achieving a median absolute error of 13 mm—representing a fivefold improvement over the unmodified baseline. Notably, approximately 25% of predictions meet the accuracy threshold required for direct robotic manipulation.

0 citationsRead paper