Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
To address inaccurate source attribution, inconsistent multilingual performance, and insufficient factual grounding in small-scale RAG models, this paper introduces Pleias-RAG-350m/1B—a lightweight, purpose-built model. Methodologically, it employs mid-scale synthetic data training, multi-stage RAG workflow modeling, cross-lingual retrieval simulation, and literal citation generation. The model natively supports verbatim citation and factual provenance tracking, integrating query routing, rewriting, and source re-ranking modules. Its key contribution is the first demonstration of consistent RAG performance across major European languages and systematic citation grounding within the sub-1B parameter regime. Experiments show that Pleias-RAG-350m/1B significantly outperforms comparable sub-4B models on benchmarks including HotPotQA and 2WikiMultihop, matching the performance of Qwen2.5-7B while enabling efficient CPU- and edge-device deployment.