MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

📅 2024-12-22
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current large language models exhibit limited capability in multi-hop reasoning tasks driven by novel or long-tail knowledge, and lack dedicated evaluation benchmarks for such settings. Method: We introduce MINTQA—the first benchmark explicitly designed to assess multi-hop question answering grounded in novel and long-tail knowledge—spanning four dimensions: question decomposition strategy, sub-question generation, retrieval-augmented generation, and dynamic decomposition-based retrieval. Our dataset comprises 28,366 high-quality QA pairs with structured sub-question chains, constructed via human curation, cross-source knowledge alignment, and joint retrieval–reasoning annotation. Contribution/Results: Evaluation across 22 state-of-the-art models reveals an average accuracy below 35%, highlighting critical bottlenecks in knowledge-aware multi-hop reasoning. MINTQA provides a reproducible, scalable, and fine-grained evaluation framework to diagnose and advance model capabilities in this domain.

Technology Category

Application Category

📝 Abstract
Large language models (LLMs) have demonstrated impressive capabilities in various reasoning tasks but face significant challenges with complex, knowledge-intensive multi-hop queries, particularly those involving new or long-tail knowledge. Existing benchmarks often fail to fully address these challenges. To bridge this gap, we introduce MINTQA (Multi-hop Question Answering on New and Tail Knowledge), a comprehensive benchmark to evaluate LLMs' capabilities in multi-hop reasoning across four critical dimensions: question handling strategy, sub-question generation, retrieval-augmented generation, and iterative or dynamic decomposition and retrieval. MINTQA comprises 10,479 question-answer pairs for evaluating new knowledge and 17,887 pairs for assessing long-tail knowledge, with each question equipped with corresponding sub-questions and answers. Our systematic evaluation of 22 state-of-the-art LLMs on MINTQA reveals significant limitations in their ability to handle complex knowledge base queries, particularly in handling new or unpopular knowledge. Our findings highlight critical challenges and offer insights for advancing multi-hop reasoning capabilities. The MINTQA benchmark is available at https://github.com/probe2/multi-hop/.
Problem

Research questions and friction points this paper is trying to address.

Complex Reasoning
Niche Knowledge
Evaluation Methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

MINTQA
Complex Reasoning
Multi-hop Questions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Edinburgh | Southeast University | University of Manchester