🤖 AI Summary
Traditional RAG systems suffer from suboptimal performance due to tight coupling among retrieval, reranking, prompt rewriting, and generation modules, hindering holistic optimization.
Method: This paper proposes the first end-to-end RAG architecture search framework, modeling the RAG configuration space as an evolvable search problem and employing genetic algorithms to jointly optimize multi-objective metrics—including recall@k/nDCG for retrieval and LLM-Judge/semantic similarity for generation. The framework encompasses nine component types across vector retrieval, reranking, and prompt rewriting.
Contribution/Results: Evaluated across six domains, the framework achieves an average 3.8% performance gain (up to +12.5% in retrieval, +7.5% in generation) while converging after exploring only 0.2% of the configuration space. It identifies robust architectural patterns transferable across datasets and quantifies how domain characteristics and question types systematically influence optimal configurations.
📝 Abstract
Retrieval-Augmented Generation (RAG) quality depends on many interacting choices across retrieval, ranking, augmentation, prompting, and generation, so optimizing modules in isolation is brittle. We introduce RAGSmith, a modular framework that treats RAG design as an end-to-end architecture search over nine technique families and 46{,}080 feasible pipeline configurations. A genetic search optimizes a scalar objective that jointly aggregates retrieval metrics (recall@k, mAP, nDCG, MRR) and generation metrics (LLM-Judge and semantic similarity). We evaluate on six Wikipedia-derived domains (Mathematics, Law, Finance, Medicine, Defense Industry, Computer Science), each with 100 questions spanning factual, interpretation, and long-answer types. RAGSmith finds configurations that consistently outperform naive RAG baseline by +3.8% on average (range +1.2% to +6.9% across domains), with gains up to +12.5% in retrieval and +7.5% in generation. The search typically explores $approx 0.2%$ of the space ($sim 100$ candidates) and discovers a robust backbone -- vector retrieval plus post-generation reflection/revision -- augmented by domain-dependent choices in expansion, reranking, augmentation, and prompt reordering; passage compression is never selected. Improvement magnitude correlates with question type, with larger gains on factual/long-answer mixes than interpretation-heavy sets. These results provide practical, domain-aware guidance for assembling effective RAG systems and demonstrate the utility of evolutionary search for full-pipeline optimization.