Institution profile

Trivago

Industry researcheurope · de
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

How Do People Quantify Naturally: Evidence from Mandarin Picture Description

Feb 10, 2026

This study investigates how speakers decide whether to quantify, how precisely to quantify, and which quantification strategies to employ during natural language production. Through a picture description task, native speakers’ spoken and written descriptions of multi-object scenes were collected under conditions with no explicit instructions, enabling the first systematic examination of quantification behavior in Mandarin Chinese in unconstrained, naturalistic settings. The findings reveal that increasing object numerosity reduces both the likelihood and precision of quantification, and that animacy interacts significantly with production modality to influence quantification strategies. The project also establishes the first corpus of Mandarin quantification based on naturally produced language, providing a foundational resource for future research in this domain.

0 citationsRead paper

Real-World Summarization: When Evaluation Reaches Its Limits

Jul 15, 2025

Assessing factual consistency of hotel highlight summaries generated by large language models (LLMs) in real-world settings remains challenging, as conventional automatic metrics perform poorly in open-domain, high-stakes commercial applications. Method: We construct a human-annotated benchmark and systematically compare lexical-overlap metrics, trainable models, and LLM-based evaluators. Contribution/Results: Simple n-gram overlap metrics—particularly ROUGE-L—exhibit strong correlation with human judgments (Spearman ρ = 0.63), significantly outperforming more complex approaches. In contrast, LLM-based evaluators suffer from prompt-induced bias, undermining annotation reliability. We propose a novel span-level error taxonomy, identifying “factual errors” and “unverifiable statements” as the highest commercial-risk error types. Our findings demonstrate that lightweight, lexically grounded fidelity assessment is not only feasible but also more robust and practically deployable than sophisticated alternatives.

0 citationsRead paper

Sentence Embeddings as an intermediate target in end-to-end summarisation

May 06, 2025

Long user reviews pose challenges for content selection in summarization, while end-to-end models suffer from insufficient coherence and information preservation under weakly aligned training corpora. Method: This paper proposes an embedding-guided extractive-abstractive summarization framework that leverages pretrained sentence embeddings (e.g., SBERT) as structured intermediate supervision—replacing conventional sentence selection probability prediction—and jointly optimizes the extractive sentence selector and abstractive sequence-to-sequence model (T5/BART) via embedding-space regression loss. Results: On a hotel review summarization dataset, our method achieves a 2.3-point ROUGE-L improvement over the state of the art; human evaluation confirms significant gains in summary relevance and fluency. The core contribution lies in introducing sentence embeddings as intermediate supervision, effectively mitigating the weak-alignment challenge inherent in long-input summarization.

0 citationsRead paper

Large Language Models as Span Annotators

Apr 11, 2025

This work addresses the high cost and low efficiency of span annotation in text, which traditionally relies on manual effort or fine-tuning encoder-based models. We propose a zero-shot/few-shot direct annotation paradigm leveraging large language models (LLMs), eliminating the need for model fine-tuning. Our method employs structured prompting combined with chain-of-thought (CoT) reasoning to elicit fine-grained, explanation-augmented span annotations from both open-source (e.g., Llama) and closed-source (e.g., GPT-series) LLMs. Key contributions include: (1) the first systematic empirical validation that LLMs can directly perform span annotation at competitive quality; (2) demonstration that reasoning-oriented LLMs achieve annotation quality, interpretability, and inter-annotator consistency (Cohen’s κ ≈ 0.4–0.6) comparable to human annotators; and (3) over 90% reduction in annotation cost versus conventional approaches. To support reproducibility and future research, we release a high-quality benchmark dataset comprising over 40,000 annotated spans.

0 citationsRead paper
Recent publications

Latest Papers

How Do People Quantify Naturally: Evidence from Mandarin Picture Description

Feb 10, 2026

This study investigates how speakers decide whether to quantify, how precisely to quantify, and which quantification strategies to employ during natural language production. Through a picture description task, native speakers’ spoken and written descriptions of multi-object scenes were collected under conditions with no explicit instructions, enabling the first systematic examination of quantification behavior in Mandarin Chinese in unconstrained, naturalistic settings. The findings reveal that increasing object numerosity reduces both the likelihood and precision of quantification, and that animacy interacts significantly with production modality to influence quantification strategies. The project also establishes the first corpus of Mandarin quantification based on naturally produced language, providing a foundational resource for future research in this domain.

0 citationsRead paper

Real-World Summarization: When Evaluation Reaches Its Limits

Jul 15, 2025

Assessing factual consistency of hotel highlight summaries generated by large language models (LLMs) in real-world settings remains challenging, as conventional automatic metrics perform poorly in open-domain, high-stakes commercial applications. Method: We construct a human-annotated benchmark and systematically compare lexical-overlap metrics, trainable models, and LLM-based evaluators. Contribution/Results: Simple n-gram overlap metrics—particularly ROUGE-L—exhibit strong correlation with human judgments (Spearman ρ = 0.63), significantly outperforming more complex approaches. In contrast, LLM-based evaluators suffer from prompt-induced bias, undermining annotation reliability. We propose a novel span-level error taxonomy, identifying “factual errors” and “unverifiable statements” as the highest commercial-risk error types. Our findings demonstrate that lightweight, lexically grounded fidelity assessment is not only feasible but also more robust and practically deployable than sophisticated alternatives.

0 citationsRead paper

Sentence Embeddings as an intermediate target in end-to-end summarisation

May 06, 2025

Long user reviews pose challenges for content selection in summarization, while end-to-end models suffer from insufficient coherence and information preservation under weakly aligned training corpora. Method: This paper proposes an embedding-guided extractive-abstractive summarization framework that leverages pretrained sentence embeddings (e.g., SBERT) as structured intermediate supervision—replacing conventional sentence selection probability prediction—and jointly optimizes the extractive sentence selector and abstractive sequence-to-sequence model (T5/BART) via embedding-space regression loss. Results: On a hotel review summarization dataset, our method achieves a 2.3-point ROUGE-L improvement over the state of the art; human evaluation confirms significant gains in summary relevance and fluency. The core contribution lies in introducing sentence embeddings as intermediate supervision, effectively mitigating the weak-alignment challenge inherent in long-input summarization.

0 citationsRead paper

Large Language Models as Span Annotators

Apr 11, 2025

This work addresses the high cost and low efficiency of span annotation in text, which traditionally relies on manual effort or fine-tuning encoder-based models. We propose a zero-shot/few-shot direct annotation paradigm leveraging large language models (LLMs), eliminating the need for model fine-tuning. Our method employs structured prompting combined with chain-of-thought (CoT) reasoning to elicit fine-grained, explanation-augmented span annotations from both open-source (e.g., Llama) and closed-source (e.g., GPT-series) LLMs. Key contributions include: (1) the first systematic empirical validation that LLMs can directly perform span annotation at competitive quality; (2) demonstration that reasoning-oriented LLMs achieve annotation quality, interpretability, and inter-annotator consistency (Cohen’s κ ≈ 0.4–0.6) comparable to human annotators; and (3) over 90% reduction in annotation cost versus conventional approaches. To support reproducibility and future research, we release a high-quality benchmark dataset comprising over 40,000 annotated spans.

0 citationsRead paper