LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对文档图像中个人身份信息(PII)泄露风险问题,通过构建LeakageBench基准数据集并评估多种检测方法来解决文档级PII去除不彻底的问题。
📝 Abstract
Real-world personally identifiable information (PII) redaction often operates on document images---scans, screenshots, and PDF renderings---where OCR errors, layout structure, and visual noise determine whether sensitive information is actually removed. Existing PII benchmarks are mostly text-centric and do not measure document-level redaction risk: a page remains unsafe if even one identifier is missed. We introduce LeakageBench, a challenge set of 500 document images with 11,954 GDPR-aligned PII annotations spanning direct identifiers, linkage keys, and contextual re-identification surfaces. We evaluate generic OCR pipelines, commercial and task-adapted OCR-dependent detectors, and OCR-free vision-language models using entity-level F1, group-wise leakage, and document-level leakage metrics. Code Interpreter raises GPT-5.5 localization F1 from 0.090 to 0.249, but critical page-level leakage remains 0.968. These results show that stronger detection and tool assistance improve localization without making most pages safe for release. LeakageBench provides a diagnostic benchmark for high-recall, spatially grounded PII redaction in document images.
Problem

Research questions and friction points this paper is trying to address.

Personally Identifiable Information
Document Images
Redaction Risk
OCR Errors
Layout Structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Document-Level Redaction
PII Detection
OCR Errors
Layout Structure
Visual Noise
🔎 Similar Papers
No similar papers found.
V
Vishnu Prasad Vijaya Kumar
Center for Artificial Intelligence and Robotics (CAIRO), Technical University of Applied Sciences Würzburg-Schweinfurt (THWS), Würzburg, Germany
S
Santhosh Venkatesh
DataX, Frankfurt am Main, Germany
Ivan P. Yamshchikov
Ivan P. Yamshchikov
Research Professor at CAIRO, THWS
natural language generationcomputational creativityempathetic aiethics of ai application