Institution profile

New Technologies for the Information Society

Industry researcheurope · hr
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Large Language Models for Summarizing Czech Historical Documents and Beyond

Aug 14, 2025

Historical document summarization for low-resource Czech has long been hindered by linguistic complexity and scarcity of annotated data. To address this, we introduce *Posel od Čerchova*, the first annotated summarization dataset for historical Czech, and achieve state-of-the-art performance on the modern Czech benchmark SumeCzech. Our approach integrates transfer learning with multilingual pretraining, adapting large language models—including Mistral and mT5—to jointly handle both modern and historical Czech texts, enabling the first unified cross-era summarization framework for the language. Key contributions are: (1) releasing the first open-source summarization dataset for historical Czech; (2) establishing the current strongest baseline for Czech summarization; and (3) empirically validating the efficacy of large language models in low-resource historical text NLP tasks, thereby providing a reusable methodological paradigm for analogous under-resourced languages.

0 citationsRead paper

Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models

Aug 14, 2025

Cross-lingual aspect-based sentiment analysis (ABSA) for low-resource languages often relies on external translation, limiting robustness and scalability—especially for complex linguistic phenomena (e.g., nested aspects, coreference). Method: This paper proposes a translation-free sequence-to-sequence framework leveraging multilingual large language models (mLLMs), introducing constrained decoding for the first time to enforce fine-grained generation control and directly output cross-lingual sentiment triplets (aspect, opinion, sentiment). Contribution/Results: The approach enables unified multilingual modeling without translation intervention and handles complex ABSA structures natively. Experiments across multiple low-resource languages show an average 10% improvement over prevailing translate-then-predict paradigms. Fine-tuned mLLMs achieve performance on par with state-of-the-art methods, whereas monolingual English LLMs underperform significantly—demonstrating both efficacy and practicality for low-resource cross-lingual ABSA.

0 citationsRead paper

Few-shot Cross-lingual Aspect-Based Sentiment Analysis with Sequence-to-Sequence Models

Aug 11, 2025

Low-resource languages suffer from limited labeled data, hindering cross-lingual aspect-based sentiment analysis (ABSA); existing approaches often rely on external translation and overlook the potential of few-shot learning in the target language. This paper proposes a lightweight few-shot transfer framework that jointly trains a sequence-to-sequence model on English data and a minimal number (10–1,000) of labeled target-language samples. Experiments across four ABSA subtasks and six low-resource languages demonstrate that just 10 target-language examples significantly outperform zero-shot baselines, while 1,000 examples surpass fully supervised monolingual models—matching the performance of sophisticated constrained decoding methods. Crucially, the approach requires no external translation tools and incurs no additional inference overhead. It establishes an efficient, general-purpose, and easily deployable paradigm for low-resource ABSA.

0 citationsRead paper

Large Language Models for Czech Aspect-Based Sentiment Analysis

Aug 11, 2025

This study presents the first systematic evaluation of 19 multilingual and Czech-specific large language models (LLMs) on Czech Aspect-Based Sentiment Analysis (ABSA), covering zero-shot, few-shot, and full fine-tuning paradigms. Under a unified evaluation framework, we analyze the impact of model scale, architecture, multilinguality, and release date on aspect term extraction and sentiment polarity classification, complemented by fine-grained error analysis. Key findings include: (1) domain-specialized smaller models significantly outperform general-purpose LLMs in zero- and few-shot settings—especially for aspect term identification; (2) fine-tuned multilingual LLMs (e.g., mT5, BLOOMZ) achieve state-of-the-art performance on Czech ABSA; and (3) model release date and multilingual capability are not decisive performance factors—domain adaptation proves more critical. These results provide empirical guidance for model selection and optimization strategies in low-resource language ABSA.

0 citationsRead paper
Recent publications

Latest Papers

Large Language Models for Summarizing Czech Historical Documents and Beyond

Aug 14, 2025

Historical document summarization for low-resource Czech has long been hindered by linguistic complexity and scarcity of annotated data. To address this, we introduce *Posel od Čerchova*, the first annotated summarization dataset for historical Czech, and achieve state-of-the-art performance on the modern Czech benchmark SumeCzech. Our approach integrates transfer learning with multilingual pretraining, adapting large language models—including Mistral and mT5—to jointly handle both modern and historical Czech texts, enabling the first unified cross-era summarization framework for the language. Key contributions are: (1) releasing the first open-source summarization dataset for historical Czech; (2) establishing the current strongest baseline for Czech summarization; and (3) empirically validating the efficacy of large language models in low-resource historical text NLP tasks, thereby providing a reusable methodological paradigm for analogous under-resourced languages.

0 citationsRead paper

Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models

Aug 14, 2025

Cross-lingual aspect-based sentiment analysis (ABSA) for low-resource languages often relies on external translation, limiting robustness and scalability—especially for complex linguistic phenomena (e.g., nested aspects, coreference). Method: This paper proposes a translation-free sequence-to-sequence framework leveraging multilingual large language models (mLLMs), introducing constrained decoding for the first time to enforce fine-grained generation control and directly output cross-lingual sentiment triplets (aspect, opinion, sentiment). Contribution/Results: The approach enables unified multilingual modeling without translation intervention and handles complex ABSA structures natively. Experiments across multiple low-resource languages show an average 10% improvement over prevailing translate-then-predict paradigms. Fine-tuned mLLMs achieve performance on par with state-of-the-art methods, whereas monolingual English LLMs underperform significantly—demonstrating both efficacy and practicality for low-resource cross-lingual ABSA.

0 citationsRead paper

Few-shot Cross-lingual Aspect-Based Sentiment Analysis with Sequence-to-Sequence Models

Aug 11, 2025

Low-resource languages suffer from limited labeled data, hindering cross-lingual aspect-based sentiment analysis (ABSA); existing approaches often rely on external translation and overlook the potential of few-shot learning in the target language. This paper proposes a lightweight few-shot transfer framework that jointly trains a sequence-to-sequence model on English data and a minimal number (10–1,000) of labeled target-language samples. Experiments across four ABSA subtasks and six low-resource languages demonstrate that just 10 target-language examples significantly outperform zero-shot baselines, while 1,000 examples surpass fully supervised monolingual models—matching the performance of sophisticated constrained decoding methods. Crucially, the approach requires no external translation tools and incurs no additional inference overhead. It establishes an efficient, general-purpose, and easily deployable paradigm for low-resource ABSA.

0 citationsRead paper

Large Language Models for Czech Aspect-Based Sentiment Analysis

Aug 11, 2025

This study presents the first systematic evaluation of 19 multilingual and Czech-specific large language models (LLMs) on Czech Aspect-Based Sentiment Analysis (ABSA), covering zero-shot, few-shot, and full fine-tuning paradigms. Under a unified evaluation framework, we analyze the impact of model scale, architecture, multilinguality, and release date on aspect term extraction and sentiment polarity classification, complemented by fine-grained error analysis. Key findings include: (1) domain-specialized smaller models significantly outperform general-purpose LLMs in zero- and few-shot settings—especially for aspect term identification; (2) fine-tuned multilingual LLMs (e.g., mT5, BLOOMZ) achieve state-of-the-art performance on Czech ABSA; and (3) model release date and multilingual capability are not decisive performance factors—domain adaptation proves more critical. These results provide empirical guidance for model selection and optimization strategies in low-resource language ABSA.

0 citationsRead paper