Projecting BrowseComp-Plus onto ClimbMix: Toward More Realistic Corpora for Agentic Search

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过将BrowseComp-Plus的问题证据迁移至更大的ClimbMix语料库中,提出了一种新的评估代理搜索的方法,解决了原语料库规模小且针对性强的问题。
📝 Abstract
The BrowseComp-Plus benchmark disentangled the evaluation of agentic search by replacing opaque web search with a fixed corpus, so that an agent's role can be separated from the retriever's. That corpus, however, holds only about 100K documents and was assembled from the supporting documents of the benchmark's own queries plus mined hard negatives, so the evidence and the distractors were both selected per query. We introduce $\text{BrowseComp-Plus}_{\text{CM}}$, which keeps the BrowseComp-Plus questions but relocates their evidence to ClimbMix, a 400B-token, 553M-document mixture of web text released by NVIDIA for pre-training language models and built without reference to any benchmark. Our main contribution is the projection pipeline that makes this possible: it decomposes each question into atomic reasoning hops and grounds every hop in the new corpus, retaining a question only when automatic verification, an independent agent, and human review all confirm that every hop is supported. The pipeline is dataset-agnostic and applies to any benchmark whose questions decompose into verifiable facts. Applied to the 830 BrowseComp-Plus test questions, our pipeline yields 57 fully grounded questions with question-level relevance judgments. Projection shifts the difficulty onto retrieval, as the strongest agent we evaluate loses five points of answer accuracy but sees its evidence recall fall from 84.3% to 21.4% while issuing 63% more search calls. As the first of a series of projections, we release the pipeline, the benchmark, and our analyses at https://github.com/castorini/cmass.
Problem

Research questions and friction points this paper is trying to address.

agentic search
corpora
realistic evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

projection pipeline
agentic search
ClimbMix corpus
retrieval difficulty
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Sahel Sharifymoghaddam
Sahel Sharifymoghaddam
University of Waterloo
Natural Language ProcessingInformation Retrieval
L
Lingwei Gu
David R. Cheriton School of Computer Science, University of Waterloo
Y
Yijun Ge
David R. Cheriton School of Computer Science, University of Waterloo
Jimmy Lin
Jimmy Lin
University of Waterloo
information retrievalnatural language processingdata managementbig data