data leakage prevention

Practices and protocols for splitting data, designing diagnostics, and implementing validation so that information does not leak between training and evaluation (or between components), ensuring realistic performance estimates and preventing confounded diagnostics.

dataleakageprevention

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.74
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$209K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical gap between compliance and effectiveness in current auditing standards—such as ASB 018—whose reliance on ambiguous language and undefined terminology obscures the potential risks associated with the use of probabilistic genotyping software in criminal justice. Through a qualitative content analysis comparing the standard’s text with five real-world audit reports, this work demonstrates for the first time that audits deemed compliant often fail to delineate the boundaries of software application. The research attributes this disconnect to structural deficiencies in the standard itself and offers concrete recommendations for revising auditing frameworks and evaluating their practical efficacy. These contributions provide both theoretical insight and actionable guidance for enhancing the governance of forensic technologies within the justice system.

AI governanceaudit standardscompliance gap

In A/B testing, rigorously evaluating novel estimation algorithms—when the true treatment effect is unobserved—remains a fundamental methodological challenge. This paper establishes, for the first time, a comprehensive theoretical framework for estimation and inference based on sample splitting: it derives the asymptotic distribution of sample-split estimators and characterizes their bias structure relative to full-sample performance; introduces a bias–variance trade-off analytical paradigm and proposes a correction-based confidence interval construction method. Leveraging statistical inference, asymptotic theory, Monte Carlo simulation, and empirical validation, the framework enables robust, production-grade evaluation of new algorithms within industrial A/B testing platforms. Theoretical results are thoroughly validated via simulation studies. The proposed infrastructure enhances A/B testing methodology by delivering an interpretable, reproducible, and deployable evaluation system.

Derives asymptotic distributions and constructs valid confidence intervalsDevelops a theoretical framework for sample splitting in A/B testingValidates results through simulations and provides implementation guidance

This study addresses the compliance challenges faced by data practitioners in machine learning systems under regulations such as the GDPR and the AI Act, particularly concerning data quality. Through semi-structured interviews with practitioners in the European Union, combined with thematic analysis of regulatory texts and engineering workflows, the research systematically uncovers a structural disconnect between regulation-driven data quality requirements and ML engineering practices. It identifies five core challenges: misalignment between legal principles and engineering implementation, fragmented data pipelines, lack of purpose-built compliance tools, ambiguous accountability, and reactive responses to audits. Building on these findings, the work proposes directions for designing compliance-oriented tooling, establishing effective governance mechanisms, and fostering cultural transformation to bridge the gap between regulatory mandates and practical ML development.

AI Actdata qualityGDPR

Current medical research agents lack domain-specific evaluation mechanisms that rigorously assess scientific validity, methodological soundness, reproducibility, and boundary safety. This work proposes MedSkillAudit—the first skill auditing framework tailored for medical research agents—which employs a hierarchical, structured pipeline to evaluate skill readiness prior to deployment. The framework incorporates expert double-blind scoring (0–100), tiered release recommendations, and high-risk flags, and quantifies agreement between the system and human experts using ICC(2,1) and weighted Cohen’s kappa. Evaluated on 75 skills, the system achieved an ICC of 0.449, surpassing inter-human rater agreement (ICC = 0.300) and demonstrating closer alignment with consensus scores (SD = 9.5 vs. 12.4), thereby validating its effectiveness and reliability.

AI governancedomain-specific evaluationmedical research agent

Current safety fine-tuning defenses are often validated by measuring the reduction in performance gaps on held-out sets; however, this metric is susceptible to sampling noise, topical artifacts, capability degradation, or non-transferable mechanisms, lacking a reliable evaluation standard. This work proposes the Acceptance Cards framework, which establishes—for the first time—a four-dimensional diagnostic criterion encompassing statistical reliability, novel semantic generalization, mechanistic alignment, and cross-task transferability, accompanied by an executable auditing toolkit for systematic validation of defense efficacy. Re-evaluating SafeLoRA on Gemma-2-2B-it across 46 experimental configurations reveals that it consistently fails to satisfy all four diagnostic criteria, exposing significant limitations in existing approaches.

defense evaluationgap reductionmechanism transfer

Latest Papers

What's happening recently
View more

This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.

clinical LLMshealthcare interoperabilityschema compliance

This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.

calibrationexchangeabilityfalse sense of safety

This study addresses the limited reliability of existing training data contamination detection methods in real-world auditing scenarios, particularly when distribution shifts occur or when reference benchmarks are substantially smaller than the pretraining corpus. Through a systematic evaluation of three dominant paradigms—LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC—the authors conduct 335 experiments across 27 open-source and state-of-the-art closed-source language models (up to 27B parameters). They identify distribution shift and small-scale benchmarks as two critical failure modes, revealing that only 199 evaluations yield correct conclusions. Current approaches suffer from high false-positive rates, low statistical power, or coarse-grained provenance resolution, rendering them inadequate for reliably verifying individual benchmark subsets and underscoring the irreplaceable value of transparent data provenance.

benchmark contaminationdata provenancedistribution shift

Current perturbation-based construct validity audits are highly sensitive to implementation details, yielding conclusions that lack transparency and reliability. This work proposes a self-audit framework that systematically identifies and formalizes five classes of audit failure modes (F1–F5). A case study encompassing two open-source instruction-tuned models and five safety benchmarks reveals that none of the audited units satisfy confirmatory criteria, exposing systemic vulnerabilities in prevailing practices. To address this, the paper introduces a six-point due diligence gating mechanism that establishes actionable standards for disclosing and retaining high-assurance audit evidence, thereby substantially enhancing the credibility and reproducibility of auditing outcomes.

AI governanceaudit failurebenchmark validity

This study addresses the widespread issue of label inaccuracies in public chest X-ray datasets, where annotations derived from radiology reports often misrepresent actual pathological findings, thereby compromising model training with unreliable supervisory signals. To mitigate this, the authors propose the Repository Supervision Auditing (RSA) framework, which employs expert image-level annotations to audit label consistency prior to model development, identify systematic biases, and construct a verified evaluation cohort. Their analysis reveals severe discrepancies between report-based labels and radiological evidence: only 1% of cases with expert-confirmed cardiomegaly were correctly labeled, while nearly half were erroneously annotated as “no abnormality.” A DenseNet121 model trained on the corrected cohort achieved a test ROC-AUC of 0.853, demonstrating that supervision auditing is a critical prerequisite for robust medical imaging AI development.

chest radiographimage-level truthmedical AI

Hot Scholars

AC

Albert Cheu

Research Scientist, Google
Differential Privacy
LW

Lili Wei

Assistant Professor at McGill University
Software EngineeringSoftware TestingSoftware AnalysisAndroid
FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse