Knowledge Synthesis Review Framework: Task-Level Benchmarking of LLM-Based Systems for Multi-Source Evidence Synthesis

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of synthesizing multi-source heterogeneous evidence—such as academic papers, reports, policies, and media content—which vary widely in quality and structure and entail high manual effort. The reliability of current large language models (LLMs) across individual synthesis subtasks remains unclear. To tackle this, the authors propose the Knowledge Synthesis Review (KSR) framework, decomposing the review process into four stages: screening, extraction, analysis, and synthesis. Using a high-agreement expert gold standard (92.2% agreement, κ=0.80), they conduct task-level evaluations of leading LLMs—including GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro—and introduce a dynamic routing mechanism that automatically selects the best-performing model under human supervision. This model-agnostic, auditable, and transparent approach significantly enhances review efficiency and coverage. Experiments reveal no single model dominates all tasks: Claude Sonnet 4 achieves the highest screening accuracy (82.8%), while GPT-5 attains the best recall (91.8%). Moreover, multi-source synthesis uncovers critical themes—such as worker well-being, small and medium enterprises, and Global South perspectives—often missed in single-source analyses.
📝 Abstract
Evidence in rapidly evolving fields is fragmented across academic studies, industry reports, policy documents, and media sources that differ in quality, structure, and purpose, making timely synthesis difficult. Large language models (LLMs) may accelerate this work, but their reliability across the distinct cognitive tasks of a review remains uncertain. We introduce the Knowledge Synthesis Review (KSR), a human-in-the-loop framework that decomposes evidence synthesis into screening, extraction, analysis, and synthesis, benchmarks LLM-based systems on each task against expert reference standards, and routes each task to the best-performing system under continuous expert validation. We evaluated GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, and NotebookLM on a 244-document benchmark subset drawn from a 1,893-document corpus on AI and work spanning four source types, against a gold standard with high inter-rater reliability (92.2% agreement, kappa = 0.80). No system led on all tasks. Claude Sonnet 4 achieved the highest screening accuracy (82.8%) and GPT-5 the highest recall (91.8%) at the expense of lower specificity. Extraction exceeded 90% agreement for titles and sources but degraded in author and reference fields. Performance declined most in interpretive analysis and cross-source synthesis, where expert judgment remained essential. A contamination check on post-cutoff documents showed no evidence that prior exposure inflated results. Applied to the full corpus, the routed workflow surfaced cross-source asymmetries and blind spots that single-source synthesis would miss, including worker well-being, small firms, and the Global South. KSR offers a transparent, auditable, model-agnostic framework for governing LLM assistance in research synthesis while preserving human accountability.
Problem

Research questions and friction points this paper is trying to address.

evidence synthesis
multi-source information
knowledge fragmentation
research synthesis
LLM reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Synthesis
Human-in-the-loop
Task-Level Benchmarking
LLM Evaluation
Evidence Synthesis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Wafa Shafqat
Wafa Shafqat
Magnet, Toronto Metropolitan University, Toronto, M5B2K3, Ontario, Canada.
M
Mark Patterson
Magnet, Toronto Metropolitan University, Toronto, M5B2K3, Ontario, Canada.; Future Skill Center, Toronto Metropolitan University, Toronto, M5B2K3, Ontario, Canada.
S
Steven N. Liss
Magnet, Toronto Metropolitan University, Toronto, M5B2K3, Ontario, Canada.; Future Skill Center, Toronto Metropolitan University, Toronto, M5B2K3, Ontario, Canada.; Office of the Vice-President Research and Innovation, Toronto Metropolitan University, Toronto, M5B2K3, Ontario, Canada.; Department of Chemistry and Biology, Toronto Metropolitan University, Toronto, M5B2K3, Ontario, Canada.; Professor Emeritus, Environmental Studies, Queen’s University, Kingston, K7L3N6, Ontario, Canada.; Professor Extraordina