Institution profile

National Library of Sweden

Academic institutioneurope · se
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

RA-ClipScore: Making Generative Model Evaluation More Interpretable

Aug 12, 2026

This work addresses the limited interpretability of existing evaluation metrics for generative models, which struggle to diagnose generation biases at both semantic attribute and spatial distribution levels. To this end, the paper proposes RA-CLIPScore, the first CLIP-based evaluation framework extended to incorporate spatial alignment. RA-CLIPScore employs a dual-prompt mechanism to disentangle competing attributes and leverages local image patch tokens to model region-wise semantic alignment. By introducing a region-specific single-attribute divergence metric, the method substantially enhances interpretability and aligns more closely with human perception. Experiments demonstrate that RA-CLIPScore exhibits greater robustness under distribution shifts or when textual prompts contain partially irrelevant attributes, and its scores show strong correlation with human judgments of visual diversity.

0 citationsRead paper

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

Aug 11, 2026

This study systematically evaluates the performance bottlenecks of speech-driven gesture generation systems across four key dimensions: motion quality, speech alignment, interactive responsiveness, and semantic expressiveness. Leveraging the Seamless Interaction dataset, the authors introduce a decoupled evaluation framework grounded in four large-scale user studies involving 869 participants and over 23,000 ratings. For the first time, they incorporate a dialogue mismatch experiment and a semantic gesture–text matching test—based on the Grounded Gestures subset—to enable independent assessment of interactivity and semantic fidelity. The results demonstrate that current systems significantly underperform human motion capture data across all evaluated dimensions, revealing critical limitations in existing approaches to generating naturalistic, contextually appropriate co-speech gestures.

0 citationsRead paper

Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

May 23, 2025

Swedish, as a medium-resource language, suffers from suboptimal performance in multilingual ASR models (e.g., Whisper) due to insufficient representation in training data. To address this, we construct the first high-diversity, cross-domain monolingual Swedish training set—built upon the largest publicly available Swedish speech corpus to date—and systematically perform supervised fine-tuning of the Whisper architecture. Our methodology incorporates rigorous data cleaning and augmentation strategies to significantly enhance model robustness. Evaluation across three major benchmarks—FLEURS, CoVo, and NST—demonstrates that our fine-tuned model achieves an average 47% relative WER reduction over Whisper-large-v3, consistently outperforming the baseline across all model sizes. This work breaks the fine-tuning performance ceiling for medium-resource languages in ASR and establishes a reproducible, high-quality monolingual data-driven paradigm for model optimization.

0 citationsRead paper
Recent publications

Latest Papers

RA-ClipScore: Making Generative Model Evaluation More Interpretable

Aug 12, 2026

This work addresses the limited interpretability of existing evaluation metrics for generative models, which struggle to diagnose generation biases at both semantic attribute and spatial distribution levels. To this end, the paper proposes RA-CLIPScore, the first CLIP-based evaluation framework extended to incorporate spatial alignment. RA-CLIPScore employs a dual-prompt mechanism to disentangle competing attributes and leverages local image patch tokens to model region-wise semantic alignment. By introducing a region-specific single-attribute divergence metric, the method substantially enhances interpretability and aligns more closely with human perception. Experiments demonstrate that RA-CLIPScore exhibits greater robustness under distribution shifts or when textual prompts contain partially irrelevant attributes, and its scores show strong correlation with human judgments of visual diversity.

0 citationsRead paper

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

Aug 11, 2026

This study systematically evaluates the performance bottlenecks of speech-driven gesture generation systems across four key dimensions: motion quality, speech alignment, interactive responsiveness, and semantic expressiveness. Leveraging the Seamless Interaction dataset, the authors introduce a decoupled evaluation framework grounded in four large-scale user studies involving 869 participants and over 23,000 ratings. For the first time, they incorporate a dialogue mismatch experiment and a semantic gesture–text matching test—based on the Grounded Gestures subset—to enable independent assessment of interactivity and semantic fidelity. The results demonstrate that current systems significantly underperform human motion capture data across all evaluated dimensions, revealing critical limitations in existing approaches to generating naturalistic, contextually appropriate co-speech gestures.

0 citationsRead paper

Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

May 23, 2025

Swedish, as a medium-resource language, suffers from suboptimal performance in multilingual ASR models (e.g., Whisper) due to insufficient representation in training data. To address this, we construct the first high-diversity, cross-domain monolingual Swedish training set—built upon the largest publicly available Swedish speech corpus to date—and systematically perform supervised fine-tuning of the Whisper architecture. Our methodology incorporates rigorous data cleaning and augmentation strategies to significantly enhance model robustness. Evaluation across three major benchmarks—FLEURS, CoVo, and NST—demonstrates that our fine-tuned model achieves an average 47% relative WER reduction over Whisper-large-v3, consistently outperforming the baseline across all model sizes. This work breaks the fine-tuning performance ceiling for medium-resource languages in ASR and establishes a reproducible, high-quality monolingual data-driven paradigm for model optimization.

0 citationsRead paper