DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense

πŸ“… 2026-07-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of detecting and explaining multi-stage Advanced Persistent Threat (APT) attacks, where existing automated defense systems often produce outputs lacking interpretability, and large language model (LLM)-generated reports frequently suffer from hallucinations or insufficient evidentiary support. To bridge this gap, the authors propose DeepFaith, a novel framework that enables evidence-driven, faithful natural language reporting for APT defense. DeepFaith integrates unified evidence representation, evidence-guided prompting, faithfulness-aware generation, and post-hoc verification to transform structured defensive outputs into natural language reports explicitly aligned with system evidence. Evaluated in real-world enterprise environments, DeepFaith improves report faithfulness from 0.68 to 0.92, reduces unsupported statements to 0.08, and achieves a temporal consistency of 0.88, substantially outperforming both template-based and LLM baselines.
πŸ“ Abstract
Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature. While recent autonomous defense systems leverage provenance graphs and learning-based models for detection and mitigation, their outputs remain largely machine-oriented and difficult for analysts to interpret. Large language models (LLMs) offer a promising interface for report generation, but often produce hallucinated or weakly grounded content. In this paper, we propose DeepFaith, an evidence-grounded framework for faithful incident reporting in multi-stage APT defense. DeepFaith transforms structured outputs from autonomous defense and explainability modules into natural-language reports that are explicitly aligned with underlying system evidence. The framework integrates a unified evidence representation, evidence-grounded prompting, faithfulness-aware generation, and post-generation verification to ensure that all generated statements are supported. Experiments in a realistic enterprise testbed demonstrate that DeepFaith improves faithfulness from 0.68 to 0.92, reduces unsupported claims from 0.32 to 0.08, and increases temporal consistency from 0.6 to 0.88, while maintaining concise reports and lower error rates than existing template-based and LLM-based solutions. These results show that evidence-grounded generation enables reliable, interpretable, and actionable reporting for security operations centers.
Problem

Research questions and friction points this paper is trying to address.

Advanced Persistent Threats
faithful reporting
evidence grounding
large language models
incident interpretation
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidence-grounded generation
faithful reporting
multi-stage APT defense
provenance-aware LLMs
post-generation verification
πŸ”Ž Similar Papers
Trung V. Phan
Trung V. Phan
Assistant Professor, Claremont Colleges (Pitzer & Scripps)
biophysicsrobophysicscondensed mattercancer chemotherapymachine learning
T
Tri Gia Nguyen
Department of Information Assurance, FPT University, Da Nang 50509, Vietnam
T
Thomas Bauschert
Chair of Communication Networks, Technische UniversitΓ€t Chemnitz, 09126 Chemnitz, Germany