Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
This study investigates whether large language models (LLMs) can automatically generate high-quality search queries and their semantic variants from information need descriptions to enable low-cost, scalable test collection construction for information retrieval (IR). We propose a prompt-engineering-based LLM query generation method and design a multidimensional evaluation framework measuring semantic similarity, document pool coverage, and relevant document overlap. To our knowledge, this is the first systematic validation of LLM-generated variants for Chinese IR test collection construction. While LLM variants exhibit slightly lower diversity than human-annotated ones, they achieve a 71.1% relevant document overlap rate at pool depth 100—significantly outperforming baseline methods. Our results demonstrate that LLMs can effectively support automated, high-fidelity test collection creation, offering a novel, practical paradigm for IR evaluation infrastructure.