audit logging

Designing tamper-evident, verifiable logs and audit trails that record actions, provenance, and policy-relevant attributes to enable post-hoc inspection and certification. This includes formats and mechanisms for attaching monotone, auditable metadata to sessions, producing human-understandable explanations, and preserving opaque evidence of origin.

auditlogging

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.66
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$194K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of traceable and tamper-resistant transparency mechanisms in large language models (LLMs) deployed in high-stakes decision-making contexts, which undermines accountability. To bridge this gap, the paper introduces the first LLM lifecycle auditing framework that integrates technical provenance with governance records. It proposes a reference architecture enabling cross-organizational traceability and implements a lightweight, open-source Python-based auditing layer. By leveraging append-only logs, event emitters, structured metadata, and an auditor interface, the system seamlessly integrates into existing LLM workflows with minimal intrusiveness. This design ensures complete, tamper-evident traceability across critical stages—including training, deployment, and monitoring—thereby facilitating robust accountability and responsibility attribution throughout the model’s lifecycle.

accountabilityaudit trailsgovernance

Current AI-assisted scientific writing lacks auditable generation processes and mechanisms for accountability, undermining the verifiability of research credibility and compliance. This work proposes a novel auditing paradigm embedded directly within the production workflow, enforcing end-to-end traceability, immutability, and third-party reproducibility of AI involvement through preregistered blind-spot indicator cards, sealed execution environments, and automated gatekeeping intercepts. Core technical components include Git-sealed lineage anchoring, hash-bound provenance tracking, red-flag interception protocols, cross-model role isolation, and programmatic assembly. In experimental validation, one project was automatically terminated when preregistered confirmatory tests triggered a No-Go decision. An open-source toolkit is released to enable independent recomputation of all core audit metrics by third parties.

AI AccountabilityAuditable AIProvenance

This work addresses the inadequacy of current AI runtime logs in providing the structured evidence necessary for legal fact-finding—such as data boundary violations or human interventions. It formalizes, for the first time, the binary factual requirements of regulatory compliance into a criterion of evidentiary sufficiency for runtime records, mandating that logs explicitly encode the legal category of events and their determinative relationships (e.g., provenance, authorization, temporal validity). By integrating legal ontologies, event-type systems, provenance semantics, and temporal validity constraints—and drawing on the law of requisite variety and the Good Regulator theorem from cybernetics—the approach exposes limitations in tamper-proof logging and generic provenance mechanisms. Validation against selected obligations of the EU AI Act demonstrates that this criterion precisely delineates the boundary between traces and hyperproperties in runtime verification, thereby establishing a verifiable foundation for compliance.

Agentic AIevidentiary adequacylegal findings

Auditable Agents

Apr 07, 2026

This study addresses the critical lack of accountability in large language model (LLM) agents following external actions, which hinders traceability and responsibility attribution. The work introduces the first systematic framework for agent auditability, articulating five core dimensions and proposing an “Auditability Card” to standardize assessment. It further develops a full-lifecycle auditing architecture integrating detection, enforcement, and recovery mechanisms. Leveraging runtime intervention, tamper-proof logging, and log-recovery techniques—validated through ecosystem-wide security evaluations and controlled experiments—the study identifies 617 security flaws across mainstream open-source LLM agent projects. Experimental results demonstrate that a pre-execution mediation layer incurs only 8.3 milliseconds of overhead and that partial reconstruction of accountability-critical information remains feasible even in the absence of complete logs.

accountabilityauditabilityevidence integrity

Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.

auditable recordsevidence validationindustrial machine learning

Latest Papers

What's happening recently
View more

Existing automated approaches for mapping cyber threat intelligence (CTI) to MITRE ATT&CK lack supporting evidence, provenance tracking, and validation history, making their credibility difficult to assess. This work proposes the first knowledge graph–driven framework for CTI governance that enables auditable management of TTP assertions through fine-grained evidence preservation, complete provenance chains, versioned trust decisions, and lossless revocation mechanisms. The framework integrates multi-extractor collaborative verification, assertion aggregation, consensus modeling, and policy-driven validation, all underpinned by versioned knowledge graph management. Evaluated on 65 CTI reports comprising 5,303 sentences, the approach achieves a precision of 90.6% under six-party consensus and efficiently supports seven categories of audit queries concerning provenance, trustworthiness, and versioning.

Cyber Threat IntelligenceMITRE ATT&CKprovenance

Current large language model (LLM) agents lack verifiability, debuggability, and auditability, and relying solely on the accuracy of final answers fails to reveal their underlying reasoning. To address this, this work proposes the first unified provenance framework for LLM agents, systematically modeling causal relationships in tool usage, memory access, and environmental interactions. It introduces a comprehensive provenance taxonomy encompassing source, granularity, representation format, and trust functions. By integrating provenance-aware representation modeling, evidence attribution, runtime safeguards, provenance-informed memory management, and trajectory observability analysis, the study shifts the evaluation paradigm from outcome correctness to process accountability. The framework consolidates existing benchmarks to define a clear pathway for process-level trustworthiness assessment and highlights key challenges, including standardized trajectory schemas, semantic-level provenance, and privacy-preserving auditing.

auditabilityevidence tracingexecution provenance

This work addresses the challenge of enabling trustworthy auditing of unstructured semantic attributes—such as code logic—while preserving the privacy of proprietary data. The authors propose a novel three-party framework that introduces a “proxy witness” paradigm, shifting verification from attested execution to attested reasoning and thereby achieving, for the first time, privacy-preserving qualitative auditing for unstructured data. By integrating trusted execution environments (TEEs), large language models (LLMs), the Model Context Protocol (MCP), and cryptographic hash chains, the framework allows verifiers to assess high-level semantic properties of private data through Boolean queries without exposing the underlying source code. Experimental results demonstrate that the approach successfully automates artifact evaluation across 21 GitHub repositories, accurately verifying five high-level semantic attributes.

confidentialityprivacy-preserving auditingproprietary data

This work addresses the challenge of ensuring trustworthy AI behavior in high-stakes, heavily regulated environments, where reliance solely on generative models, output safeguards, or post-hoc audits proves insufficient to prevent unacceptable execution trajectories. To this end, the paper introduces the Proposal–Certification–Execution (PCE) framework, which formalizes trajectory permissibility as an explicit safety property requiring prior certification. Central to PCE are the Permissibility Machine and a verifiable certificate mechanism that enforce a “no certificate, no execution” policy, thereby providing pre-execution trust guarantees. Integrating a policy system Π, a language for executable trajectories, proof-carrying execution, and privacy-preserving techniques, the framework establishes a structured pre-execution certification process and advances a new evaluation paradigm centered on certifiably permissible trajectories—shifting trustworthy AI from output monitoring toward pre-execution verification.

certified tracesexecution controlpermissibility

Existing autonomous commercial protocols struggle to achieve interoperable, tamper-proof auditing and event temporal verification across heterogeneous domains. This work proposes a verifiable global event timeline architecture that constructs a reproducible, tamper-resistant AI fraud intelligence training pipeline by formalizing event schemas, employing deterministic batching, leveraging Merkle append-only commitments, and anchoring events to blockchain-based timestamps. The approach innovatively integrates cryptographic fraud markers—binding risk labels with anchored evidence—and a data provenance model to establish a verifiable, traceable, AI-ready intelligence layer. Evaluated on a prototype processing 50,000 events, the system constructs Merkle trees in just 47 milliseconds, achieves end-to-end verification in under 0.013 milliseconds, and exhibits logarithmic proof size growth, yielding a 14.4× improvement in verification efficiency over linear scanning.

agentic commercefraud intelligencetamper-evident auditability

Hot Scholars

FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
YH

Yintong Huo

Singapore Management University
AI4SEAIOpsLog analysisMLLM for SE
BI

Brittany I. Davidson

Associate Professor; University of Bath
behavioural analyticscomputational social sciencebehavioural sciencehuman-computer interaction
YL

Yang Liu

Nanyang Technological University
AgentSoftware EngineeringCyber SecurityTrustworthy AI