One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI
本文研究了多预行动控制下AI系统的动作校正问题,提出了一种再校正协议以恢复每项动作的合理性,并探讨了证据替换和资源预算调整的不同顺序对控制平面语义的影响。
本文研究了多预行动控制下AI系统的动作校正问题,提出了一种再校正协议以恢复每项动作的合理性,并探讨了证据替换和资源预算调整的不同顺序对控制平面语义的影响。
This work addresses silent decision errors in AI agents caused by metadata defects—such as stale prices or superseded records—that evade model detection and conventional data quality alerts. The authors propose a runtime data quality gating mechanism that identifies such defects prior to action execution and integrates downstream-only repair strategies to prevent high-cost errors stemming from unreliable evidence. Innovatively treating evidence integrity as a system-level dimension orthogonal to model capability, they design a model-free oracle based on task-decision geometry to analyze defect propagation patterns. Experimental results demonstrate full recovery of performance loss within the method’s coverage scope, with the oracle achieving a mean absolute error of 0.015 (Pearson r = 0.876), empirically validating the “flat-staircase” phenomenon wherein defect impact remains independent of model proficiency.
This work addresses the absence of real-time cost and carbon emission constraints in existing AI agent systems, which typically rely on post-hoc monitoring. The authors propose the first verifiable, zero-overrun budget and carbon gating mechanism embedded directly within the agent execution loop, grounded in the SARC governance framework and enforcing predictive governance at four critical decision points. Their approach integrates quantile conformal calibration, soft Lagrangian penalties, and both synthetic and real task arrival models, and they release a complete open-source governance library. Experiments on real-world multi-step tasks demonstrate zero budget overruns and achieve end-to-end reductions of 47–55% in token usage, monetary cost, and carbon emissions, with all results fully reproducible.
Current AI agent systems lack effective runtime governance over tool usage and multi-agent behaviors, leading to a disconnect between compliance requirements and actual execution. This work proposes the SARC framework, which formalizes governance constraints as executable specifications—comprising source, category, predicate, and validation point—and embeds them as first-class objects within the agent’s operational loop. SARC enforces these constraints through a four-tier mechanism: pre-action gating, in-action monitoring, post-action auditing, and escalation routing, enabling real-time verification and enforcement before, during, and after actions, as well as along escalation pathways. In procurement task evaluations, SARC achieved zero violations of hard constraints and reduced soft constraint exceedances by 89.5%, with residual violations primarily attributable to execution stack errors rather than environmental risks, thereby demonstrating its effectiveness and traceability.
This study addresses the risk of non-price vertical foreclosure in the commercialization of generative AI, particularly within inference, distribution, and routing layers, which may undermine market competition. It presents the first formal model capturing mechanisms such as quality-of-service discrimination in AI inference markets, routing bias at the assistant layer, and tiered access restrictions. Integrating game-theoretic analysis, Logit demand functions, and symmetric competition assumptions, the model is calibrated using projected 2026 data from four major providers. The analysis identifies equilibrium conditions for service-quality gaps and their boundary relationships with key market parameters. The paper proposes a “neutral inference” regulatory framework, with quantitative results indicating that Google and OpenAI possess the strongest foreclosure capabilities, while Microsoft’s multi-model strategy constrains its leverage. Implementation of the proposed framework could generate annual net welfare gains amounting to tens of billions of dollars.
本文研究了多预行动控制下AI系统的动作校正问题,提出了一种再校正协议以恢复每项动作的合理性,并探讨了证据替换和资源预算调整的不同顺序对控制平面语义的影响。
This work addresses silent decision errors in AI agents caused by metadata defects—such as stale prices or superseded records—that evade model detection and conventional data quality alerts. The authors propose a runtime data quality gating mechanism that identifies such defects prior to action execution and integrates downstream-only repair strategies to prevent high-cost errors stemming from unreliable evidence. Innovatively treating evidence integrity as a system-level dimension orthogonal to model capability, they design a model-free oracle based on task-decision geometry to analyze defect propagation patterns. Experimental results demonstrate full recovery of performance loss within the method’s coverage scope, with the oracle achieving a mean absolute error of 0.015 (Pearson r = 0.876), empirically validating the “flat-staircase” phenomenon wherein defect impact remains independent of model proficiency.
This work addresses the absence of real-time cost and carbon emission constraints in existing AI agent systems, which typically rely on post-hoc monitoring. The authors propose the first verifiable, zero-overrun budget and carbon gating mechanism embedded directly within the agent execution loop, grounded in the SARC governance framework and enforcing predictive governance at four critical decision points. Their approach integrates quantile conformal calibration, soft Lagrangian penalties, and both synthetic and real task arrival models, and they release a complete open-source governance library. Experiments on real-world multi-step tasks demonstrate zero budget overruns and achieve end-to-end reductions of 47–55% in token usage, monetary cost, and carbon emissions, with all results fully reproducible.
Current AI agent systems lack effective runtime governance over tool usage and multi-agent behaviors, leading to a disconnect between compliance requirements and actual execution. This work proposes the SARC framework, which formalizes governance constraints as executable specifications—comprising source, category, predicate, and validation point—and embeds them as first-class objects within the agent’s operational loop. SARC enforces these constraints through a four-tier mechanism: pre-action gating, in-action monitoring, post-action auditing, and escalation routing, enabling real-time verification and enforcement before, during, and after actions, as well as along escalation pathways. In procurement task evaluations, SARC achieved zero violations of hard constraints and reduced soft constraint exceedances by 89.5%, with residual violations primarily attributable to execution stack errors rather than environmental risks, thereby demonstrating its effectiveness and traceability.
This study addresses the risk of non-price vertical foreclosure in the commercialization of generative AI, particularly within inference, distribution, and routing layers, which may undermine market competition. It presents the first formal model capturing mechanisms such as quality-of-service discrimination in AI inference markets, routing bias at the assistant layer, and tiered access restrictions. Integrating game-theoretic analysis, Logit demand functions, and symmetric competition assumptions, the model is calibrated using projected 2026 data from four major providers. The analysis identifies equilibrium conditions for service-quality gaps and their boundary relationships with key market parameters. The paper proposes a “neutral inference” regulatory framework, with quantitative results indicating that Google and OpenAI possess the strongest foreclosure capabilities, while Microsoft’s multi-model strategy constrains its leverage. Implementation of the proposed framework could generate annual net welfare gains amounting to tens of billions of dollars.