No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM评估中人名使用问题,提出PUN协议构建并验证未知但形式合理的姓名,结合Wikidata、网络筛选和搜索重验证方法。
📝 Abstract
Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence, and no ambiguity signals under a documented validation run, and introduce PUN (Plausible Unknown Names), a protocol for constructing and validating such names, combining Wikidata-derived components, web-enabled LLM screening, and controlled search revalidation. We report acceptance rate, reproducibility, ablations, and a 204-participant human study, finding accepted names are more name-like than controls while participants recover person evidence in only 3% of cases. We release 300 names with comparison controls.
Problem

Research questions and friction points this paper is trying to address.

person names
LLM evaluation
evidential status
memorisation
retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Plausible Unknown Names
LLM Evaluation
Wikidata-derived components
web-enabled LLM screening
🔎 Similar Papers
No similar papers found.