RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决推荐系统中错误请求识别问题,提出RPCBench基准,评估大语言模型对错误前提的检测、诊断及处理能力。
📝 Abstract
Large language models are increasingly used as interactive recommender assistants. Their evaluation should therefore go beyond plausible item recommendation and test whether they can recognize flawed recommendation requests. Existing recommender benchmarks mainly assess ranking, generation, or preference satisfaction, while existing error-detection benchmarks are usually not grounded in recommendation-specific user and candidate evidence. To address this gap, we introduce RPCBench, a benchmark for evaluating Recommender-Premise Critique: the ability to detect, diagnose, and properly handle faulty premises in natural-language recommendation requests. RPCBench contains evidence-grounded test instances from five recommendation domains and covers ten types of premise failures. Each instance provides a visible recommendation context and a corrupted user query. We further design a fine-grained evaluation framework that measures proactive detection, error localization, post-detection handling strategy, and evidence faithfulness. Through a systematic evaluation of 11 LLMs, we find that proactive detection is the main bottleneck in Recommender-Premise Critique, and models perform worst on underspecified-premise errors. We also observe that target-critical information density matters more than redundant evidence, and that longer reasoning does not monotonically improve critique quality: performance peaks at intermediate reasoning length, while overly long reasoning is accompanied by an overthinking penalty. The code is available at https://github.com/ZhongruChen/RPCBench.
Problem

Research questions and friction points this paper is trying to address.

Recommender-Premise Critique
faulty premises
natural-language recommendation requests
error detection
interactive recommender assistants
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recommender-Premise Critique
proactive detection
evidence faithfulness
premise failures
reasoning length
🔎 Similar Papers
Z
Zhongru Chen
School of Artificial Intelligence, Jilin University
Y
Yuan Wu
School of Artificial Intelligence, Jilin University
Yi Chang
Yi Chang
Jilin University
Information RetrievalData MiningNatural Language ProcessingMachine LearningArtificial Intelligence