SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses silent decision errors in AI agents caused by metadata defects—such as stale prices or superseded records—that evade model detection and conventional data quality alerts. The authors propose a runtime data quality gating mechanism that identifies such defects prior to action execution and integrates downstream-only repair strategies to prevent high-cost errors stemming from unreliable evidence. Innovatively treating evidence integrity as a system-level dimension orthogonal to model capability, they design a model-free oracle based on task-decision geometry to analyze defect propagation patterns. Experimental results demonstrate full recovery of performance loss within the method’s coverage scope, with the oracle achieving a mean absolute error of 0.015 (Pearson r = 0.876), empirically validating the “flat-staircase” phenomenon wherein defect impact remains independent of model proficiency.
📝 Abstract
Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payload and betrayed only by freshness, lineage, or provenance. Such a defect never enters the agent's context, and an agent cannot doubt data it cannot see. On a priced replenishment benchmark, a competent agent silently converts an injected metadata-borne defect into a costly action about 60% of the time, with zero data-quality flags and behavioral doubt markers at chance (AUC <= 0.50). Across four model tiers spanning roughly 15x in inference price, the rate stays flat: capability does not buy skepticism. A metadata-aware pre-action gate with downstream-only remediation recovers the loss fully on the signals its predicates cover and not at all on those they miss. A model-free oracle derived from the task's decision geometry tracks the measured rates with MAE 0.015 (Pearson r = 0.876, interval coverage 15/16 cells), giving the flat ladder an analytical form. Evidence integrity is a systems axis distinct from model capability; mitigation depends on enforcement placement and predicate coverage. Code, frozen results, and a deterministic analysis pipeline: https://github.com/besanson/dqSarc
Problem

Research questions and friction points this paper is trying to address.

data-quality gating
metadata-borne defects
agentic AI
evidence integrity
silent evidence defects
Innovation

Methods, ideas, or system contributions that make the work stand out.

data-quality gating
metadata-borne defects
agentic AI
downstream-only remediation
evidence integrity
💼 Related Jobs
No related jobs found.
G
Gaston Besanson
Universidad Torcuato Di Tella