Assessing Reliability of BERT-Based Models on Question Answering Tasks

πŸ“… 2026-08-11
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of systematic evaluation of reliability in question-answering models by proposing a novel framework that integrates Monte Carlo Dropout (MCD) with input semantic paraphrasing perturbations, requiring no modification to the standard inference pipeline. The authors conduct a comprehensive analysis of BERT, RoBERTa, ALBERT, and DistilBERT on the SQuAD and QuAC datasets. Experimental results demonstrate that MCD effectively captures prediction stability, with RoBERTa exhibiting the highest reliability, while ALBERT and DistilBERT show notably lower stability. This work not only validates MCD as a robust metric for model reliability but also establishes a new paradigm for evaluating the robustness of question-answering systems.
πŸ“ Abstract
Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical applications. Recent advancements in natural language processing (NLP), particularly those based on transformer architectures, have significantly accelerated progress across various NLP tasks. This study focuses on the reliability of transformer-based question answering (QA) models, specifically BERT models and its variants (RoBERTa, ALBERT, DistilBERT). These encoder-only pretrained transformers have demonstrated remarkable accuracy in QA tasks that can be treated as classification tasks. However, their reliability remains underexplored. This study evaluates the reliability of four BERT-based models by assessing response stability under two conditions: (1) internal model variations induced via Monte Carlo Dropout (MCD) and (2) input perturbations through paraphrasing. Using the SQuAD and QuAC datasets, we investigate how dropout rates affect prediction consistency and whether lexical changes impact answer stability. Our findings reveal that RoBERTa maintains higher reliability, whereas AlBERT and DistilBERT exhibit significant inconsistencies. Statistical analyses confirm that enabling MCD during prediction does not disrupt inference dynamics, validating its effectiveness as a reliability metric. These findings underscore the importance of evaluating both accuracy and stability in QA models to ensure stability in real-world applications.
Problem

Research questions and friction points this paper is trying to address.

reliability
question answering
BERT
model stability
input perturbation
Innovation

Methods, ideas, or system contributions that make the work stand out.

reliability estimation
Monte Carlo Dropout
input perturbation
BERT variants
question answering stability
πŸ”Ž Similar Papers
No similar papers found.
P
Pooja Yadav
Department of Mathematics, Malaviya National Institute of Technology, Jaipur, Rajasthan, India
P
Priyanka Harjule
Department of Mathematics, Malaviya National Institute of Technology, Jaipur, Rajasthan, India
Basant Agarwal
Basant Agarwal
Central University of Rajasthan
Deep LearningNatural Language ProcessingMachine LearningSentiment AnalysisSentic Computing
M
Marko Robnik Ε ikonja
University of Ljubljana, Faculty of Computer and Information Science, Ljubljana, Slovenia