incident response planning

Designs incident response plans and playbooks for responsible AI and security incidents, specifying detection, escalation, communication, and remediation procedures.

incidentresponseplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$197K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional expert-manual-based cybersecurity response methods, which struggle to adapt to dynamic attack scenarios and evolving recovery objectives, as well as the instability of existing large-model approaches in long-horizon tasks. The authors propose an end-to-end agent planning framework that innovatively models event states using a graph structure (Graph-as-State), incorporates a phase-aware agent routing mechanism, and establishes a verifiable experience reuse loop to guide action selection and state updates. The system integrates multi-agent large language models with experience retrieval augmentation and execution feedback verification, enabling dynamic, stable, and evolvable response planning within a Docker-based network range simulation environment. Experimental results demonstrate that the proposed method achieves a normalized defense score of 0.94 across 100 simulated scenarios, representing a 9.5% improvement over the strongest baseline.

adaptive responseagentic planningcyberattack recovery

Current cybersecurity incident response (IR) playbooks and intrusion models are developed independently and asynchronously, leading to a fundamental semantic misalignment between threat modeling and response execution at the model level. Method: This paper proposes a unified modeling framework built upon the Security Modelling Framework (SMF), introducing the first Sequential AND Attack Tree formalism that enables precise semantic alignment between attack steps and response actions; develops an automated tool for transforming attack trees into structured IR playbooks; and constructs the first publicly available, cross-domain intrusion-response joint model suite—compatible across nine critical infrastructure sectors, including energy and transportation. Contribution/Results: The framework significantly enhances the operationalizability of intrusion models, achieves deep coupling between threat modeling and response orchestration, and uncovers novel, actionable mappings between attack paths and response actions—enabling synergistic analysis and improved decision support.

Enhancing cyber attack analysis using Security Modelling FrameworkImproving critical infrastructure security through unified threat modelingIntegrating intrusion models with incident response playbooks for consistency

This work addresses the inefficiency of traditional manual security incident response and the limitations of existing automated approaches, which are either difficult to deploy or suffer from unreliable planning due to hallucinations when relying solely on large language models (LLMs). To overcome these challenges, the paper proposes a novel multi-scale intelligent response architecture that integrates decision-theoretic planning with a lightweight LLM. The framework uniquely combines digital twins, a tactical-level rollout planner, and an operational-level LLM agent, establishing a dual-scale (tactical–operational) coordination mechanism to enable reliable and executable automated responses in simulated environments. Experimental results across three attack scenarios demonstrate that the proposed approach reduces average recovery time by 15.1% and improves success rate by 33.6% compared to state-of-the-art LLM-based baselines.

automated planningdecision-theoretic planningdigital twin

This study addresses the absence of actionable international standards for determining when AI incidents warrant escalation from national to cross-border coordinated responses. It proposes a systematic, multi-jurisdictional escalation framework that integrates eight assessment criteria, gated decision points, and threshold mechanisms to balance local policy flexibility with global coordination. Through regulatory analysis (e.g., SB 53, EU AI Act), cross-sectoral response framework comparisons, structured case testing, and flowchart modeling, the research identifies three design patterns in developer-led reporting systems that contribute to underreporting and highlights how ambiguous definitions and data gaps critically undermine detection efficacy. Validation across ten real-world and variant incidents demonstrates the framework’s practical utility while exposing significant deficiencies in current regimes regarding timeliness and operational feasibility.

AI incident escalationescalation criteriainternational coordination

Designing Incident Reporting Systems for Harms from General-Purpose AI

Nov 08, 2025
KW
Kevin Wei
🏛️ RAND Centre for the Governance of AI

In response to escalating safety and rights risks posed by general-purpose artificial intelligence (GPAI), this paper proposes the first systematic reporting framework for GPAI incidents. Drawing on a systematic literature review and cross-case analysis of high-stakes domains—including aviation and healthcare—as well as regulatory practices in the U.S. and EU, the study identifies seven core dimensions: policy objectives, reporting entities, incident typologies, reporting modalities (mandatory vs. voluntary), near-miss inclusion, anonymity safeguards, and legal immunity provisions. It critically examines the trade-offs among safety learning, cross-organizational information sharing, and legal interoperability inherent in each mechanism. The resulting framework offers policymakers and researchers an actionable, theory-informed blueprint for designing GPAI incident reporting infrastructure—addressing a critical gap in GPAI risk governance and advancing the institutional foundations for responsible AI development and deployment.

Addressing safety and rights harms from general-purpose AI adoptionDeveloping institutional frameworks for AI incident reporting systemsEstablishing design dimensions for effective incident reporting processes

Latest Papers

What's happening recently
View more

This work addresses a critical gap in the safety mechanisms of large language model (LLM) agents, which predominantly focus on preventive measures and lack capabilities for responding to and recovering from incidents after they occur. To bridge this gap, the authors propose AIR—the first incident response framework tailored for LLM agents—that integrates semantic anomaly detection, tool-driven containment, and recovery mechanisms within the agent’s execution loop. AIR also leverages a domain-specific language to automatically generate protective rules, enabling autonomous management across the full incident lifecycle. Experimental evaluation demonstrates that AIR achieves over 90% success rates in detection, remediation, and eradication across three representative agent types, with automatically generated rules performing comparably to handcrafted ones, all while maintaining acceptable runtime overhead.

autonomous systemsfailure recoveryincident response

Enterprise incident response is often delayed due to reliance on static playbooks and manual analysis. This work proposes the first supervised automated response system that integrates a multi-agent architecture with security ontologies—namely MITRE ATT&CK, D3FEND, and NIST CSF 2.0—to enable structured, low-risk, and auditable response workflows. The system employs role-based agents, a Planner-Validator closed-loop for action verification, and a Moderator gateway for security oversight. It incorporates an action catalog with risk scoring and maintains an append-only audit log for full traceability. Evaluated on a test set of 120 incidents, the system improves the FP-aware IRS F1 score from 0.61 to 0.84 and reduces harmful actions to 0.0%, substantially outperforming static baselines.

alert-to-containment delayanalyst-driven triageenterprise intrusion response

This work addresses the challenges of frequent, rapidly propagating, and complex failures in hyperscale cloud networks, which are difficult to manage efficiently through traditional manual operations. The authors propose a multi-agent collaborative architecture with progressive autonomy, featuring hierarchical agent design, skill-driven standardized tool invocation, structured encoding of operational knowledge, layered security controls, and closed-loop validation mechanisms to enable automated fault detection, diagnosis, and remediation. Deployed in production environments of major cloud providers, the system achieves over 90% autonomous resolution rates for common failure categories, significantly enhancing operational autonomy while ensuring safety and reliability.

AI AgentsAutonomous Incident ResolutionHyperscale Networks

Current AI incident governance frameworks lack consistency in defining, categorizing, monitoring, and reporting incidents, which constrains the depth and accuracy of post-deployment failure analysis. This study addresses this gap through a systematic literature review and comparative analysis across multiple governance frameworks, thereby identifying and synthesizing key inconsistencies that span existing mechanisms. The work reveals systemic deficiencies in data collection practices, classification logics, and analytical rigor, and elucidates critical misalignments among core governance components. By clarifying these structural disconnects, the research establishes a theoretical foundation and proposes a coordinated pathway toward a unified, standardized framework for AI incident governance.

AI incident governanceclassificationdefinitions

This work addresses inefficiencies in Network Operations Centers (NOCs)—including fragmented information retrieval, verbose ticketing, and loss of contextual continuity during handoffs—stemming from data silos. To mitigate these challenges, the authors propose ORBIT, an intelligent agent system integrated into the ServiceNow platform. ORBIT employs a modular, layered architecture that encapsulates task logic into versioned, testable “skills,” ensuring reliable and scalable operation within constrained behavioral boundaries. Its core components comprise a centralized reasoning engine, an MCP protocol interface to ESnet, a semantic search layer, an operational chat interface, and a LiteLLM model gateway. Evaluated on six initial tasks and rapidly adapted to two new ones, ORBIT significantly reduces operational steps, eliminates known error patterns, and sees broad adoption of its reusable components, thereby lowering cognitive load and accelerating incident response.

cognitive loadcontext lossincident resolution

Hot Scholars

KH

Kim Hammar

Postdoc at the University of Melbourne
Distributed SystemsReinforcement LearningCyber SecurityDecision Theory
MR

Milena Radenkovic

University of Nottingham UK, Microsoft Research Ltd, Cambridge, UK
Complex networksComplex GraphsAI and ML and AnalyticsSecurity and Privacy
TL

Tao Li

City University of Hong Kong
Game TheoryReinforcement LearningSecurityIntelligent Transportation
SS

Sebastian Schinzel

Münster University of Applied Sciences, Fraunhofer SIT, Athene
Computer Security
TN

Timur Naushirvanov

PhD Student
Network and Data ScienceInternational DevelopmentEducational Policy and Innovations