Score
Designing disclosure and audit practices that clearly document when and how generative AI is used, and structuring audits and reports so their scope, limitations, and governance implications are transparent and actionable for stakeholders.
Current AI governance research lacks systematic integration of diverse frameworks and practices, with notable gaps in the operationalizability of key mechanisms and the implementation of inclusive, stakeholder-centered approaches. To address this, we conduct a rapid three-tier literature review, systematically synthesizing nine authoritative IEEE/ACM reviews published between 2020 and 2024. We introduce the novel “thematic semantic synthesis” analytical paradigm to identify high-frequency governance frameworks (e.g., the EU AI Act, NIST AI Risk Management Framework), core principles (e.g., transparency, accountability), and stakeholder role distributions. Our analysis reveals four critical knowledge gaps in AI governance scholarship and practice. Based on these findings, we propose a rigorously grounded, organizationally feasible governance roadmap—bridging theoretical advancement and real-world implementation. This work contributes both empirical evidence and methodological innovation to advance AI governance research and practice.
This paper identifies a core dilemma in organizational responsible AI governance: ambiguous responsibility boundaries across AI lifecycle stages and a lack of role- and stage-appropriate operational tools. Methodologically, the study systematically reviews over 220 responsible AI tools and proposes a novel two-dimensional (Actor, Stage) classification framework, integrating systematic review, meta-analysis, and qualitative coding. It identifies three critical governance gaps: (1) unclear accountability attribution, (2) absence of empirical validation for most tools, and (3) severe coverage imbalance across actors and stages. Results show that >80% of tools target developers during data and modeling phases; tools for leadership, deployers, end users, and stages such as value proposition definition and deployment are virtually absent. Moreover, >90% of tools lack empirical evidence. The study establishes a theoretically grounded, empirically benchmarked framework to advance actor–stage–aligned AI governance tool ecosystems.
AI systems frequently lack auditability—their capacity to undergo independent, lifecycle-wide evaluation against ethical, legal, and technical standards—due to fragmented regulatory frameworks and the absence of systematic auditing practices. Method: This study proposes an “audit-embedded governance” framework that proactively integrates auditing requirements into AI development workflows. It establishes a holistic methodology comprising standardized compliance documentation, dynamic risk assessment protocols, multi-tiered governance structures, and interoperable auditing tools. The approach emphasizes cross-stakeholder collaboration and socio-technical co-design, while advocating for international regulatory alignment and mutual recognition. Contribution/Results: The framework clarifies actionable pathways for scaling AI audits across organizational and jurisdictional boundaries. It delivers a practical, implementation-oriented guide for building trustworthy, transparent, and accountable AI governance systems, alongside concrete policy recommendations for regulators, developers, and auditors.
Ordinary users struggle to effectively participate in auditing generative AI content, and their feedback is rarely incorporated by industry practitioners. Method: We propose WeAudit—the first structured framework supporting user-centered participatory AI auditing—integrating reflective interaction design and a cross-role feedback loop to enable both individual and collaborative auditing, thereby aligning user-identified issues, problem representation, and developer responses. Grounded in human-AI collaboration principles, WeAudit was validated through qualitative formal research, iterative prototyping, a three-week in-the-wild user study, and in-depth interviews with AI practitioners. Contribution/Results: WeAudit significantly enhances users’ ability to detect potential AI harms and generates audit reports that are both engineering-interpretable and actionable. These outputs received strong endorsement from AI practitioners, demonstrating the framework’s practical viability and impact on closing the gap between end-user insights and industrial AI development practices.
Current AI auditing tools predominantly focus on model performance evaluation, failing to support end-to-end accountability practices—including harm identification, evidence construction, stakeholder engagement, and intervention advocacy. Method: We conducted in-depth interviews with 35 practitioners and systematically crawled, cataloged, and coded 435 auditing tools to develop a需求–capability mapping framework and an ecosystem gap diagnostic model. Contribution/Results: Our analysis reveals systematic deficiencies across four critical capabilities: traceability, multi-stakeholder participation, evidentiary chain generation, and intervention support—constituting the first empirical diagnosis of such gaps. Building on these findings, we propose the “AI Accountability Infrastructure” paradigm, shifting beyond narrow assessment-centric design toward cross-stage coordination and multi-role adaptability. The study delivers an empirically grounded, prioritized roadmap for designing next-generation, accountability-oriented AI auditing tools.
The lack of functional transparency in AI systems impedes stakeholders’ ability to assess their fairness and accuracy. Method: This study proposes a “Compliance-by-Design” transparency framework that uniquely embeds legal compliance requirements *a priori* into the explainable AI (XAI) design process, integrating interdisciplinary perspectives from human–computer interaction, algorithmic ethics, legal compliance, and socio-technical systems analysis. Contribution/Results: The framework establishes a functionally transparent system that is perceptible, interactive, and accountable to users. It yields a reusable set of transparency design guidelines and a multi-dimensional evaluation pathway. Empirical validation demonstrates significant improvements in AI systems’ comprehensibility, auditability, and societal trustworthiness in real-world deployments. By systematically aligning technical design with normative values and regulatory expectations, this work provides a methodological foundation for developing fair, trustworthy, and value-aligned AI systems.
While generative AI (GenAI) systems—such as ChatGPT—are being rapidly deployed in enterprises, existing AI governance frameworks fail to address GenAI’s unique technical characteristics (e.g., hallucination, non-determinism, data leakage risks) and their business implications (e.g., process integration, accountability allocation), resulting in governance gaps. Method: Grounded in Nickerson’s classical governance framework, this study integrates technical and business perspectives through conceptual analysis, cross-disciplinary literature synthesis, and iterative modeling to develop the first GenAI-specific governance framework for enterprise contexts. Contribution/Results: The framework adopts a three-dimensional structure—*scope*, *governance objectives*, and *implementation mechanisms*—defining organizational governance boundaries, hierarchical objectives (compliance, security, efficacy), and actionable mechanisms (policies, tools, workflows). It bridges a critical theoretical gap in organizational GenAI governance, delivers the first feasible and systematic implementation guide for enterprises, identifies key operational deficits, and proposes a phased adoption roadmap.
This study addresses a critical gap between compliance and effectiveness in current auditing standards—such as ASB 018—whose reliance on ambiguous language and undefined terminology obscures the potential risks associated with the use of probabilistic genotyping software in criminal justice. Through a qualitative content analysis comparing the standard’s text with five real-world audit reports, this work demonstrates for the first time that audits deemed compliant often fail to delineate the boundaries of software application. The research attributes this disconnect to structural deficiencies in the standard itself and offers concrete recommendations for revising auditing frameworks and evaluating their practical efficacy. These contributions provide both theoretical insight and actionable guidance for enhancing the governance of forensic technologies within the justice system.
Current binary declarations of generative AI use fail to accurately capture students’ nuanced engagement across diverse academic tasks, thereby hindering both academic integrity and the development of AI literacy. This work proposes a domain-specific disclosure framework for computer science education, integrating a generative AI usage taxonomy with educational assessment theory to address writing and programming assignments. The framework guides students—according to their cognitive developmental stage—to transparently report AI involvement in specific phases such as writing planning, content generation, and code refinement. Moving beyond the simplistic “used/not used” dichotomy, this structured, task-oriented declaration mechanism effectively distinguishes legitimate assistance from academic misconduct, fosters student metacognition, and lays the groundwork for institutions to implement honest assessment practices and future workplace standards for AI documentation.
Public sector actors increasingly rely on vendor-provided AI transparency documents, such as FactSheets, for accountability and risk assessment, yet their practical utility remains empirically unexamined. This study addresses this gap through semi-structured interviews and a systematic content analysis of FactSheets published by the GovAI Coalition, revealing for the first time that these documents function dually as both marketing instruments and disclosure mechanisms in practice. The findings indicate that while FactSheets alone are insufficient to support robust technical evaluation, they play a critical role in fostering trust, enabling stakeholder alignment, and sustaining ongoing governance dialogues. Building on these insights, the paper proposes reconceptualizing FactSheets not merely as static informational artifacts but as relational governance tools that facilitate dynamic, iterative engagement between public institutions and AI vendors.
Current evaluations of AI systems predominantly rely on static benchmarks, which fail to capture behavioral risks in dynamic real-world environments. This work formalizes AI auditing as an uncertainty-aware, dynamic constraint monitoring problem across the system’s entire lifecycle, targeting critical attributes such as fairness and safety while integrating sociotechnical norms with statistical risk control. By developing a theoretical framework and supporting infrastructure for continuous auditing, the study advances AI governance beyond one-off testing toward ongoing, reliable, and accountable oversight mechanisms.
This work addresses the lack of governance support—particularly for multi-party collaboration, controlled access, and traceable workflows—in existing AI experimentation environments. The authors propose and implement a governance-aware, multi-tenant AI sandbox platform based on a layered reference architecture that decouples presentation, control, execution, and data management layers. Integrated approval workflows and audit logging mechanisms structurally capture experimental context and governance decisions. The platform enables controlled onboarding, cross-project collaboration, and compliance verification, and for the first time facilitates the generation of reusable evaluation evidence, thereby enhancing experiment comparability and auditability. Its effectiveness has been validated in industrial–academic collaborative settings, yielding key insights into deployment and evolutionary practices for such governance-integrated platforms.