Evaluating Modern RAG: Textual, Multimodal, Dense, and Late Interaction Pipelines

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种数据驱动的选择方法,帮助选择适合文档语料库的RAG管道,通过评估文本和多模态管道来平衡检索性能与系统效率。
📝 Abstract
Retrieval-augmented generation (RAG) systems have traditionally relied on text-based pipelines that extract and retrieve information from documents. While efficient and lightweight, these approaches often struggle with documents where meaning is conveyed through layout, tables, and visual elements. Recent advances in multimodal pipelines, powered by vision-language models (VLMs), improve retrieval quality by jointly encoding visual and textual signals, but at increased computational and memory cost. We propose a quantitative, data-driven selection methodology that guides practitioners in choosing the most appropriate RAG pipeline for a given document corpus based on empirical effectiveness and resource constraints. We evaluate contemporary textual and multimodal pipelines, including dense and late-interaction architectures, analyze their trade-offs, and provide actionable guidance for balancing retrieval performance with system efficiency.
Problem

Research questions and friction points this paper is trying to address.

retrieval-augmented generation
multimodal pipelines
vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

data-driven selection
multimodal pipelines
retrieval-augmented generation (RAG)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Emre Kuru
Özyeğin University, Istanbul, Türkiye
M
Mehmet Onur Keskin
Özyeğin University, Istanbul, Türkiye