Institution profile

Nuremberg University of Technology

Academic institutioneurope · de
Official website
Research library151linked papers
Opportunities0open roles
Selected work

Representative Papers

Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation

Aug 17, 2025Interspeech

This work addresses the lack of structured hierarchy in speech-to-text transcripts by proposing a multilevel topic segmentation method that integrates prosodic pause features with LoRA-finetuned large language models to automatically generate hierarchical outlines comprising topics and subtopics. It introduces LoRA-based fine-tuning for the first time to the task of multilevel transcript segmentation, designs a unified evaluation metric tailored to hierarchical structure, and enhances boundary detection accuracy through the incorporation of speech pause information. The approach significantly outperforms existing baselines on English meeting corpora as well as Portuguese and German lecture datasets, demonstrating strong effectiveness and cross-lingual generalization across diverse scenarios.

2 citationsRead paper

From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications

Jan 12, 2026arXiv.org

This study addresses the inefficiency in psychological triage caused by vague email subject lines in online mental health counseling. It proposes a hierarchical evaluation framework to classify and rank within-category six-word German subject lines generated by eleven large language models, integrating assessments from both professional counselors and AI scorers to systematically evaluate their practical utility in mental health contexts. The work presents the first comparative analysis of closed-source and open-source models on German-language mental health tasks, examining performance and ethical trade-offs, and employs Krippendorff’s α and Spearman’s ρ to quantify inter-rater agreement and score correlations. Results demonstrate that German-specific fine-tuning substantially enhances model performance, offering empirical support for deploying AI systems in mental health settings that balance privacy and efficacy.

1 citations1 influentialRead paper

Privacy-Aware Visual Language Models

May 27, 2024arXiv.org

Vision-language models (VLMs) exhibit poor recognition accuracy for visual privacy content—such as passports and fingerprints—and existing evaluation datasets suffer from inconsistent labeling, hindering rigorous privacy-safety assessment and optimization. Method: We introduce PrivBench, the first benchmark dedicated to visual privacy understanding, and construct PrivTune, a lightweight instruction-tuning dataset. Leveraging models like TinyLLaVa and MiniGPT-v2, we propose a privacy-aware few-shot instruction-tuning paradigm that preserves general vision-language capabilities (e.g., VQA) while enhancing privacy-sensitive image recognition. Contribution/Results: Our approach achieves state-of-the-art performance on PrivBench—surpassing GPT-4V—without degrading general-purpose capabilities. This work establishes the first systematic evaluation standard and efficient adaptation framework for visual privacy in VLMs, providing foundational support for privacy-aware VLM research.

1 citationsRead paper

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Aug 13, 2026

This work addresses a critical gap in the evaluation of vision-language models (VLMs), which typically emphasize perceptual and reasoning accuracy while neglecting behavioral reliability under missing or misleading visual evidence. To this end, we introduce SciFigBench—a challenging diagnostic benchmark for scientific figure understanding comprising over 34,000 test samples—designed with image perturbations, adversarial probes, and selective blurring, complemented by human-annotated labels and multidimensional metrics including MQM scores and reasoning accuracy. We further propose the A-R-I framework to systematically assess whether models Acknowledge insufficient evidence, Resist misleading cues, and Infer cautiously under uncertainty. Empirical results reveal that while GPT-5.2 achieves high accuracy, it frequently hallucinates; in contrast, Gemini 3.1 Pro demonstrates comparable performance with markedly higher reliability, explicitly acknowledging uncertainty in 71% of cases and attaining a resistance-to-misleading score of 0.91.

0 citationsRead paper

CPDA: Class-Conditional Path Distribution Alignment for Unsupervised Time-Series Domain Adaptation

Aug 10, 2026

This work addresses distribution shifts in unsupervised time series domain adaptation caused by variations across users, sensors, or environments. The authors propose a non-adversarial framework that aligns class-conditional path distributions between source and target domains in a latent space, leveraging both ground-truth labels and soft pseudo-labels to preserve class semantics during alignment. The key innovation lies in the first-time introduction of class-conditional path distribution alignment, enabled by a composite kernel function integrating semantic features, temporal structure, frequency-domain information, and low-rank path signatures. Theoretical risk bounds are provided to support the approach. Using this signature-spectral kernel as a discrepancy measure with CNN, ResNet18, and TCN backbones, the method significantly outperforms 30 existing techniques—including those based on discrepancy minimization, adversarial learning, and pseudo-labeling—across 13 benchmarks, demonstrating robust and state-of-the-art cross-domain classification performance.

0 citationsRead paper
Recent publications

Latest Papers

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Aug 13, 2026

This work addresses a critical gap in the evaluation of vision-language models (VLMs), which typically emphasize perceptual and reasoning accuracy while neglecting behavioral reliability under missing or misleading visual evidence. To this end, we introduce SciFigBench—a challenging diagnostic benchmark for scientific figure understanding comprising over 34,000 test samples—designed with image perturbations, adversarial probes, and selective blurring, complemented by human-annotated labels and multidimensional metrics including MQM scores and reasoning accuracy. We further propose the A-R-I framework to systematically assess whether models Acknowledge insufficient evidence, Resist misleading cues, and Infer cautiously under uncertainty. Empirical results reveal that while GPT-5.2 achieves high accuracy, it frequently hallucinates; in contrast, Gemini 3.1 Pro demonstrates comparable performance with markedly higher reliability, explicitly acknowledging uncertainty in 71% of cases and attaining a resistance-to-misleading score of 0.91.

0 citationsRead paper

CPDA: Class-Conditional Path Distribution Alignment for Unsupervised Time-Series Domain Adaptation

Aug 10, 2026

This work addresses distribution shifts in unsupervised time series domain adaptation caused by variations across users, sensors, or environments. The authors propose a non-adversarial framework that aligns class-conditional path distributions between source and target domains in a latent space, leveraging both ground-truth labels and soft pseudo-labels to preserve class semantics during alignment. The key innovation lies in the first-time introduction of class-conditional path distribution alignment, enabled by a composite kernel function integrating semantic features, temporal structure, frequency-domain information, and low-rank path signatures. Theoretical risk bounds are provided to support the approach. Using this signature-spectral kernel as a discrepancy measure with CNN, ResNet18, and TCN backbones, the method significantly outperforms 30 existing techniques—including those based on discrepancy minimization, adversarial learning, and pseudo-labeling—across 13 benchmarks, demonstrating robust and state-of-the-art cross-domain classification performance.

0 citationsRead paper

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Aug 10, 2026

This study addresses the limitations of current machine translation evaluation practices, which are predominantly English-centric, overlook regional cultural differences, and are susceptible to data contamination, thereby failing to assess model robustness on localized content. To remedy this, the authors propose a source-language contrastive evaluation paradigm and introduce Cultivar—a benchmark derived from a localized subset of FLORES—that compares model performance on localized versus non-localized translations to detect data contamination and evaluate regional adaptability. This framework extends the unit of evaluation from language pairs to localized content, systematically uncovering performance disparities across regional contexts. Experiments on 32 open-source models reveal insufficient robustness in specialized translation systems, evidence of overfitting to FLORES in some cases, and a consistent bias favoring U.S.-localized content.

0 citationsRead paper

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Aug 10, 2026

This study addresses the gap in existing large language model (LLM) evaluation frameworks, which often neglect public administration values and Dutch linguistic characteristics. Through expert consultation, user surveys, and civil servant interviews, the authors develop “Grip on LLMs,” a systematic evaluation framework tailored for Dutch government applications. It assesses over thirty general-purpose and Dutch-specific models across six dimensions: factuality, honesty, social bias, energy consumption, cost, and training data transparency. Innovatively integrating public governance principles with local language requirements, the work introduces a multi-dimensional trade-off perspective, revealing that factuality and honesty are governed by distinct mechanisms. A visualization tool is also developed to support non-technical decision-makers in model selection. Findings indicate no single model dominates across all criteria; high performance frequently entails greater environmental and economic costs, and bias shows no significant correlation with model capability. The results are publicly released as an accessible, multi-stakeholder model overview platform.

0 citationsRead paper

Fourier Self-Supervision for Fine-Grained Generalized Category Discovery

Aug 09, 2026

This work addresses the vulnerability of existing generalized category discovery methods to superficial visual cues, which hinders their ability to capture intrinsic attributes of fine-grained novel categories. To overcome this limitation, the authors propose a self-supervised dual-frequency filtering mechanism based on Fourier transform: low-pass filtering extracts high-level semantic content, while high-pass filtering preserves edge and texture details. These components are processed separately in distinct latent spaces and subsequently fused to construct a more robust and comprehensive feature representation. Notably, this is the first approach to incorporate frequency-domain information into generalized category discovery. Without requiring prior knowledge of the number of unknown classes, the method significantly enhances fine-grained novel category recognition and achieves state-of-the-art performance across multiple benchmark datasets.

0 citationsRead paper