Institution profile

PleIAs

Research institution
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

Apr 25, 2025

To address inaccurate source attribution, inconsistent multilingual performance, and insufficient factual grounding in small-scale RAG models, this paper introduces Pleias-RAG-350m/1B—a lightweight, purpose-built model. Methodologically, it employs mid-scale synthetic data training, multi-stage RAG workflow modeling, cross-lingual retrieval simulation, and literal citation generation. The model natively supports verbatim citation and factual provenance tracking, integrating query routing, rewriting, and source re-ranking modules. Its key contribution is the first demonstration of consistent RAG performance across major European languages and systematic citation grounding within the sub-1B parameter regime. Experiments show that Pleias-RAG-350m/1B significantly outperforms comparable sub-4B models on benchmarks including HotPotQA and 2WikiMultihop, matching the performance of Qwen2.5-7B while enabling efficient CPU- and edge-device deployment.

0 citationsRead paper

What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

Apr 10, 2025

This paper identifies severe construct validity deficiencies in the HellaSwag benchmark—including grammatical errors, misleading prompts, ambiguous answer options, and spurious statistical shortcuts—that undermine its reliability for evaluating language models’ commonsense reasoning. Through ablation studies across multiple LLM scales, answer-text isolation tests, controlled “Lorem ipsum” prompt baselines, and fine-grained human annotation, we quantitatively demonstrate for the first time that over 65% of model predictions are context-agnostic, rendering evaluations highly susceptible to superficial surface patterns. We propose six essential design principles for next-generation commonsense reasoning benchmarks and open-source GoldenSwag, a rigorously curated subset addressing these flaws. Experiments show that GoldenSwag substantially improves both reliability (inter-annotator agreement) and validity (correlation with human judgment and robustness to distractors), enabling more trustworthy model selection and capability attribution.

0 citationsRead paper
Recent publications

Latest Papers

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

Apr 25, 2025

To address inaccurate source attribution, inconsistent multilingual performance, and insufficient factual grounding in small-scale RAG models, this paper introduces Pleias-RAG-350m/1B—a lightweight, purpose-built model. Methodologically, it employs mid-scale synthetic data training, multi-stage RAG workflow modeling, cross-lingual retrieval simulation, and literal citation generation. The model natively supports verbatim citation and factual provenance tracking, integrating query routing, rewriting, and source re-ranking modules. Its key contribution is the first demonstration of consistent RAG performance across major European languages and systematic citation grounding within the sub-1B parameter regime. Experiments show that Pleias-RAG-350m/1B significantly outperforms comparable sub-4B models on benchmarks including HotPotQA and 2WikiMultihop, matching the performance of Qwen2.5-7B while enabling efficient CPU- and edge-device deployment.

0 citationsRead paper

What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

Apr 10, 2025

This paper identifies severe construct validity deficiencies in the HellaSwag benchmark—including grammatical errors, misleading prompts, ambiguous answer options, and spurious statistical shortcuts—that undermine its reliability for evaluating language models’ commonsense reasoning. Through ablation studies across multiple LLM scales, answer-text isolation tests, controlled “Lorem ipsum” prompt baselines, and fine-grained human annotation, we quantitatively demonstrate for the first time that over 65% of model predictions are context-agnostic, rendering evaluations highly susceptible to superficial surface patterns. We propose six essential design principles for next-generation commonsense reasoning benchmarks and open-source GoldenSwag, a rigorously curated subset addressing these flaws. Experiments show that GoldenSwag substantially improves both reliability (inter-annotator agreement) and validity (correlation with human judgment and robustness to distractors), enabling more trustworthy model selection and capability attribution.

0 citationsRead paper