Benchmarking Large Language Models for Quebec Insurance: From Closed-Book to Retrieval-Augmented Generation

📅 2026-03-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge faced by Quebec insurance consumers in comprehending complex financial contracts due to a lack of professional guidance. To this end, the authors construct AEPC-QA, a private benchmark comprising 807 multiple-choice questions, and systematically evaluate the legal accuracy of 51 large language models under two paradigms: closed-book generation and retrieval-augmented generation (RAG). The findings reveal that chain-of-thought reasoning substantially enhances performance; RAG acts as a knowledge equalizer, boosting accuracy by over 35 percentage points for weaker models, though it may also degrade performance due to contextual interference. Notably, general-purpose large models outperform specialized fine-tuned smaller models, illustrating a “specialization paradox.” The top-performing model achieves 79% accuracy, approaching human expert levels.

Technology Category

Application Category

📝 Abstract
The digitization of insurance distribution in the Canadian province of Quebec, accelerated by legislative changes such as Bill 141, has created a significant"advice gap", leaving consumers to interpret complex financial contracts without professional guidance. While Large Language Models (LLMs) offer a scalable solution for automated advisory services, their deployment in high-stakes domains hinges on strict legal accuracy and trustworthiness. In this paper, we address this challenge by introducing AEPC-QA, a private gold-standard benchmark of 807 multiple-choice questions derived from official regulatory certification (paper) handbooks. We conduct a comprehensive evaluation of 51 LLMs across two paradigms: closed-book generation and retrieval-augmented generation (RAG) using a specialized corpus of Quebec insurance documents. Our results reveal three critical insights: 1) the supremacy of inference-time reasoning, where models leveraging chain-of-thought processing (e.g. o3-2025-04-16, o1-2024-12-17) significantly outperform standard instruction-tuned models; 2) RAG acts as a knowledge equalizer, boosting the accuracy of models with weak parametric knowledge by over 35 percentage points, yet paradoxically causing"context distraction"in others, leading to catastrophic performance regressions; and 3) a"specialization paradox", where massive generalist models consistently outperform smaller, domain-specific French fine-tuned ones. These findings suggest that while current architectures approach expert-level proficiency (~79%), the instability introduced by external context retrieval necessitates rigorous robustness calibration before autonomous deployment is viable.
Problem

Research questions and friction points this paper is trying to address.

insurance advice gap
legal accuracy
trustworthiness
Quebec insurance
automated advisory
Innovation

Methods, ideas, or system contributions that make the work stand out.

retrieval-augmented generation
chain-of-thought reasoning
legal accuracy
context distraction
specialization paradox
David Beauchemin
David Beauchemin
Laval University
Machine learningNatural Language ProcessingDeep LearningLawInsurance
R
Richard Khoury
Group for Research in Artificial Intelligence of Laval University (GRAIL), Université Laval, Quebec City, Quebec, Canada