Institution profile

National Library of Norway

Academic institutioneurope · no
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Stringalign: Moving beyond summary statistics with a transparent Unicode-aware tool for evaluating automatic transcription models

Jun 14, 2026

This work addresses the inconsistency and poor reproducibility of character error rate (CER) and word error rate (WER) metrics in evaluating automatic transcription models, which stem from opaque preprocessing and ambiguous definitions of characters and words in existing text alignment tools. To resolve this, the authors propose Stringalign, a lightweight Python library that introduces FAIR principles to transcription evaluation for the first time. Stringalign enables reproducible, fine-grained error analysis through Unicode-aware transparent normalization, flexible tokenization strategies, and character- and word-level alignment algorithms. Coupled with interactive visualizations, Stringalign significantly enhances evaluation consistency and interpretability across OCR, handwritten text recognition (HTR), and automatic speech recognition (ASR) tasks, thereby facilitating effective model diagnosis and selection.

0 citationsRead paper

Grammaticality Judgments in Humans and Language Models: Revisiting Generative Grammar with LLMs

Dec 11, 2025

This study investigates whether large language models (LLMs), trained solely on surface-level sequences, can spontaneously acquire hierarchical syntactic sensitivity—specifically, whether they exhibit structure-dependent generalization without explicit grammatical supervision. Method: Focusing on two canonical structure-dependent phenomena—subject–auxiliary inversion and parasitic gap licensing—the authors employ controlled prompt engineering to elicit grammaticality judgments from models including GPT-4 and LLaMA-3, then rigorously assess structural generalization via systematic minimal-pair comparisons. Contribution/Results: Results demonstrate that LLMs consistently distinguish grammatical from ungrammatical sentences across novel constructions, exhibiting functional syntactic competence that transcends linear sequential cues. Critically, this hierarchical structural sensitivity emerges robustly despite the absence of overt syntactic annotations or architectural biases toward hierarchy. This work provides the first systematic empirical evidence that LLMs spontaneously develop hierarchical syntactic representations from surface input alone—challenging the generative grammar tenet that syntactic structure must be innately encoded.

0 citationsRead paper

BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications

Sep 29, 2025

Spanish legal texts—particularly BOE (Boletín Oficial del Estado) decrees and notices—lack accessible, concise summaries, exacerbating information overload for non-expert readers. Method: We introduce BOE-XSUM, the first extreme summarization dataset for Spanish official legal documents, comprising 3,648 human-written plain-language summaries. We fine-tune medium-scale models—including BERTIN and GPT-J 6B—in a supervised setting and compare them against zero-shot baselines. Summary accuracy is evaluated using exact-match metrics. Contribution/Results: BOE-XSUM fills a critical gap in Spanish legal extreme summarization. Fine-tuned models substantially outperform zero-shot generation, achieving a best-case accuracy of 41.6%—a 24-percentage-point improvement—demonstrating that domain-specific data coupled with lightweight fine-tuning significantly enhances the generation of comprehensible, legally faithful summaries.

0 citationsRead paper
Recent publications

Latest Papers

Stringalign: Moving beyond summary statistics with a transparent Unicode-aware tool for evaluating automatic transcription models

Jun 14, 2026

This work addresses the inconsistency and poor reproducibility of character error rate (CER) and word error rate (WER) metrics in evaluating automatic transcription models, which stem from opaque preprocessing and ambiguous definitions of characters and words in existing text alignment tools. To resolve this, the authors propose Stringalign, a lightweight Python library that introduces FAIR principles to transcription evaluation for the first time. Stringalign enables reproducible, fine-grained error analysis through Unicode-aware transparent normalization, flexible tokenization strategies, and character- and word-level alignment algorithms. Coupled with interactive visualizations, Stringalign significantly enhances evaluation consistency and interpretability across OCR, handwritten text recognition (HTR), and automatic speech recognition (ASR) tasks, thereby facilitating effective model diagnosis and selection.

0 citationsRead paper

Grammaticality Judgments in Humans and Language Models: Revisiting Generative Grammar with LLMs

Dec 11, 2025

This study investigates whether large language models (LLMs), trained solely on surface-level sequences, can spontaneously acquire hierarchical syntactic sensitivity—specifically, whether they exhibit structure-dependent generalization without explicit grammatical supervision. Method: Focusing on two canonical structure-dependent phenomena—subject–auxiliary inversion and parasitic gap licensing—the authors employ controlled prompt engineering to elicit grammaticality judgments from models including GPT-4 and LLaMA-3, then rigorously assess structural generalization via systematic minimal-pair comparisons. Contribution/Results: Results demonstrate that LLMs consistently distinguish grammatical from ungrammatical sentences across novel constructions, exhibiting functional syntactic competence that transcends linear sequential cues. Critically, this hierarchical structural sensitivity emerges robustly despite the absence of overt syntactic annotations or architectural biases toward hierarchy. This work provides the first systematic empirical evidence that LLMs spontaneously develop hierarchical syntactic representations from surface input alone—challenging the generative grammar tenet that syntactic structure must be innately encoded.

0 citationsRead paper

BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications

Sep 29, 2025

Spanish legal texts—particularly BOE (Boletín Oficial del Estado) decrees and notices—lack accessible, concise summaries, exacerbating information overload for non-expert readers. Method: We introduce BOE-XSUM, the first extreme summarization dataset for Spanish official legal documents, comprising 3,648 human-written plain-language summaries. We fine-tune medium-scale models—including BERTIN and GPT-J 6B—in a supervised setting and compare them against zero-shot baselines. Summary accuracy is evaluated using exact-match metrics. Contribution/Results: BOE-XSUM fills a critical gap in Spanish legal extreme summarization. Fine-tuned models substantially outperform zero-shot generation, achieving a best-case accuracy of 41.6%—a 24-percentage-point improvement—demonstrating that domain-specific data coupled with lightweight fine-tuning significantly enhances the generation of comprehensible, legally faithful summaries.

0 citationsRead paper