🤖 AI Summary
This study addresses the limitations of traditional pipeline-based architectures in Web research, which struggle to meet the dynamic and generative demands of the large language model (LLM) era. It systematically investigates how generative paradigms—particularly retrieval-augmented generation (RAG)—are reshaping core Web research tasks such as information retrieval, question answering, and recommendation systems, while also extending to emerging applications like Web summarization and educational tools. As the first comprehensive survey of generative retrieval’s impact on Web research, this work delineates the evolutionary trajectory from conventional pipelines to LLM-driven approaches, identifies key challenges, and proposes actionable directions for optimization, thereby offering both theoretical grounding and practical guidance for the development of next-generation Web intelligence systems.
📝 Abstract
Web research and practices have evolved significantly over time, offering users diverse and accessible solutions across a wide range of tasks. While advanced concepts such as Web 4.0 have emerged from mature technologies, the introduction of large language models (LLMs) has profoundly influenced both the field and its applications. This wave of LLMs has permeated science and technology so deeply that no area remains untouched. Consequently, LLMs are reshaping web research and development, transforming traditional pipelines into generative solutions for tasks like information retrieval, question answering, recommendation systems, and web analytics. They have also enabled new applications such as web-based summarization and educational tools. This survey explores recent advances in the impact of LLMs-particularly through the use of retrieval-augmented generation (RAG)-on web research and industry. It discusses key developments, open challenges, and future directions for enhancing web solutions with LLMs.