Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

๐Ÿ“… 2026-04-30
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆๅ‡บไบ†ไธ€็งๆ–ฐ็š„ๅŸบๅ‡†ๆต‹่ฏ•ๆ–นๆณ•๏ผŒ็”จไบŽ่ฏ„ไผฐๅคงๅž‹่ฏญ่จ€ๆจกๅž‹ๅœจ็Ÿฅ่ฏ†่พน็•ŒไธŠ็š„่กจ็Žฐ๏ผŒ้€š่ฟ‡ๅŒบๅˆ†ๆœ‰ๆ”ฏๆŒ็š„ๅ›ž็ญ”ไธŽๆ— ๆ”ฏๆŒ็š„็Œœๆต‹๏ผŒๅนถ่€ƒ่™‘ๆ•ฐๆฎๆฑกๆŸ“็ญ‰ๅ› ็ด ใ€‚
๐Ÿ“ Abstract
Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Knowledge-Boundary Evaluation
Data Contamination
Prompt Idiosyncrasy
Answerability
Innovation

Methods, ideas, or system contributions that make the work stand out.

contamination-aware
multi-zone benchmark
knowledge-boundary evaluation
abstention
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
R
Renwei Meng
Anhui University, Hefei, China
B
Bowen Zhang
Anhui University, Hefei, China
J
Jian Wang
Guangxi University, Nanning, China
X
Xican Wang
Anhui University, Hefei, China
Haoyi Wu
Haoyi Wu
ShanghaiTech University
X
Xuanyan Qiu
Anhui University, Hefei, China
Shengan Yang
Shengan Yang
University of Cambridge
RoboticsControl