FRESCO: Benchmarking and Optimizing Re-rankers for Evolving Semantic Conflict in Retrieval-Augmented Generation

๐Ÿ“… 2026-04-14
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge that existing rerankers, when evaluated under static benchmarks, struggle to handle temporally evolving factual inconsistencies in retrieval-augmented generation and often favor semantically rich yet outdated documents. To this end, the authors introduce FRESCO, a novel benchmark that constructs time-sensitive queryโ€“document pairs using historical Wikipedia snapshots, enabling systematic evaluation of rerankers in dynamic information environments. They further propose a multi-task instruction fine-tuning framework that leverages Pareto-optimal strategies to balance performance between evolving and non-evolving knowledge tasks. Experimental results demonstrate that the proposed approach achieves up to a 27% performance gain on evolving-knowledge tasks while maintaining competitive results on static tasks, significantly enhancing model sensitivity to factual timeliness.

Technology Category

Application Category

๐Ÿ“ Abstract
Retrieval-Augmented Generation (RAG) is a key approach to mitigating the temporal staleness of large language models (LLMs) by grounding responses in up-to-date evidence. Within the RAG pipeline, re-rankers play a pivotal role in selecting the most useful documents from retrieved candidates. However, existing benchmarks predominantly evaluate re-rankers in static settings and do not adequately assess performance under evolving information -- a critical gap, as real-world systems often must choose among temporally different pieces of evidence. To address this limitation, we introduce FRESCO (Factual Recency and Evolving Semantic COnflict), a benchmark for evaluating re-rankers in temporally dynamic contexts. By pairing recency-seeking queries with historical Wikipedia revisions, FRESCO tests whether re-rankers can prioritize factually recent evidence while maintaining semantic relevance. Our evaluation reveals a consistent failure mode across existing re-rankers: a strong bias toward older, semantically rich documents, even when they are factually obsolete. We further investigate an instruction optimization framework to mitigate this issue. By identifying Pareto-optimal instructions that balance Evolving and Non-Evolving Knowledge tasks, we obtain gains of up to 27% on Evolving Knowledge tasks while maintaining competitive performance on Non-Evolving Knowledge tasks.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
re-ranker
temporal dynamics
semantic conflict
factual recency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation
re-ranker benchmarking
temporal dynamics
instruction optimization
semantic conflict
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Sohyun An
Sohyun An
UCLA
SearchOptimizationGenerative Models
Hayeon Lee
Hayeon Lee
Research Scientist, Meta GenAI
neural architecture searchmeta-learningAutoML
S
Shuibenyang Yuan
Meta Superintelligence Labs
C
Chun-cheng Jason Chen
Meta Superintelligence Labs
Cho-Jui Hsieh
Cho-Jui Hsieh
University of California, Los Angeles
Machine LearningOptimization
V
Vijai Mohan
Meta Superintelligence Labs
A
Alexander Min
Meta Superintelligence Labs