An Auditable Pipeline for Fuzzy Full-Text Screening in Systematic Reviews: Integrating Contrastive Semantic Highlighting and LLM Judgment

📅 2025-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Systematic reviews suffer from low full-text screening efficiency, primarily because critical evidence is dispersed across heterogeneous long documents, making static binary inclusion/exclusion rules inadequate. Method: We propose an auditable fuzzy screening pipeline that integrates contrastive semantic embeddings with large language model (LLM)-based judgment, supporting multi-label, progressive eligibility assessment. For the first time, we combine Mamdani-type fuzzy logic control with dynamic thresholding and confidence decay to enable fine-grained, interpretable, multi-criteria decision-making. The pipeline incorporates domain-adapted embeddings, contrastive cosine similarity, semantic highlighting, and LLM-based judgment modules. Results: On a fully positive dataset, recall per eligibility criterion reaches 75.0%–87.5%; strict adherence to all criteria achieves 50.0% inclusion rate; average screening time per document drops from 20 minutes to under 1 minute; and human-AI agreement reaches 96.1%. Our approach significantly improves recall, traceability, and screening efficiency.

Technology Category

Application Category

📝 Abstract
Full-text screening is the major bottleneck of systematic reviews (SRs), as decisive evidence is dispersed across long, heterogeneous documents and rarely admits static, binary rules. We present a scalable, auditable pipeline that reframes inclusion/exclusion as a fuzzy decision problem and benchmark it against statistical and crisp baselines in the context of the Population Health Modelling Consensus Reporting Network for noncommunicable diseases (POPCORN). Articles are parsed into overlapping chunks and embedded with a domain-adapted model; for each criterion (Population, Intervention, Outcome, Study Approach), we compute contrastive similarity (inclusion-exclusion cosine) and a vagueness margin, which a Mamdani fuzzy controller maps into graded inclusion degrees with dynamic thresholds in a multi-label setting. A large language model (LLM) judge adjudicates highlighted spans with tertiary labels, confidence scores, and criterion-referenced rationales; when evidence is insufficient, fuzzy membership is attenuated rather than excluded. In a pilot on an all-positive gold set (16 full texts; 3,208 chunks), the fuzzy system achieved recall of 81.3% (Population), 87.5% (Intervention), 87.5% (Outcome), and 75.0% (Study Approach), surpassing statistical (56.3-75.0%) and crisp baselines (43.8-81.3%). Strict "all-criteria" inclusion was reached for 50.0% of articles, compared to 25.0% and 12.5% under the baselines. Cross-model agreement on justifications was 98.3%, human-machine agreement 96.1%, and a pilot review showed 91% inter-rater agreement (kappa = 0.82), with screening time reduced from about 20 minutes to under 1 minute per article at significantly lower cost. These results show that fuzzy logic with contrastive highlighting and LLM adjudication yields high recall, stable rationale, and end-to-end traceability.
Problem

Research questions and friction points this paper is trying to address.

Addressing fuzzy decision-making in systematic review full-text screening
Integrating contrastive semantic highlighting with LLM judgment
Improving recall and efficiency in multi-criteria article inclusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fuzzy decision pipeline with contrastive semantic highlighting
LLM adjudication with confidence scores and rationales
Dynamic threshold mapping using Mamdani fuzzy controller
P
P. Mortezaagha
Ottawa Hospital Research Institute, Ottawa, Ontario, Canada
A
A. Rahgozar
Ottawa Hospital Research Institute, Ottawa, Ontario, Canada