Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了检测器定义的数据集中沉默污染问题,通过贝叶斯方法分析不同数据池中的真实与虚假事件比例,揭示了污染对估计精度的影响。
📝 Abstract
Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component (residual 1.1e-16), with phantoms outnumbering true events 473 to 308, so contamination is a second signal with detector-inherited shape, not additive noise. Contamination direction depends on the estimator: on identical windows one statistic is diluted and another inflated because its denominator is also contaminated. A common normalization turns the estimator into a mean of ratios whose expectation need not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.
Problem

Research questions and friction points this paper is trying to address.

ML datasets
detector-defined datasets
silent contamination
prevalence
Bayes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayes
dataset precision
phantom contamination
detector-defined dataset
true-positive prevalence
J
Jia Huang
Guanghua School of Management, Peking University
Y
Yankai Wan
College of Artificial Intelligence, Jilin University
Y
Yangjun Ou
School of Mathematical Sciences, Peking University