Institution profile

Splunk Inc.

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Observability for Delegated Execution in Agentic AI Systems

Jun 08, 2026

This work addresses the challenge that standard observability data in large language model agent systems often fails to distinguish execution traces across different delegation assignments, rendering cross-tool and cross-system delegation behaviors untraceable. To overcome this limitation, the paper proposes an agent-aware observability infrastructure that leverages a lightweight gateway and a unified information model to bind delegation context at runtime. This approach enables, for the first time, precise reconstruction of delegation scopes without relying on heuristic time windows. By supporting fine-grained behavioral forensics and direct forensic queries, the method significantly enhances observability and auditability of delegated executions in heterogeneous multi-agent systems.

0 citationsRead paper

Policy-Guided Threat Hunting: An LLM enabled Framework with Splunk SOC Triage

Mar 25, 2026

This work proposes an automated, dynamic threat-hunting framework that integrates AI agents with the Splunk SIEM platform to address the evolving challenges posed by advanced persistent threats (APTs) and the inefficiencies in processing massive volumes of heterogeneous logs within security operations centers (SOCs). The framework uniquely combines a reconstruction-based autoencoder, a two-layer deep reinforcement learning architecture, and a large language model to establish a policy-guided, context-aware autonomous hunting mechanism. This enables an end-to-end closed-loop pipeline—from traffic ingestion and anomaly detection to risk prioritization. Experimental results demonstrate that the approach adaptively aligns with diverse SOC objectives and effectively identifies suspicious and malicious network traffic on both public and simulated datasets, significantly enhancing analysts’ decision-making efficiency.

0 citationsRead paper

VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

Jul 11, 2025

Large-scale labels automatically generated by multimodal foundation models (e.g., CLIP, LLaVA) lack ground-truth annotations; existing evaluation methods rely on limited metrics or small-sample inspections, hindering detection of latent errors—especially in open-vocabulary image segmentation. Method: We propose the first human-in-the-loop visual analytics framework tailored for this task. It integrates multimodal output analysis, visual clustering-based diagnosis, interactive label correction, and an expert feedback loop to enable fine-grained quality assessment and iterative refinement. Contribution/Results: By introducing visual analytics into the quality assurance pipeline for auto-generated labels, our approach overcomes the limitations of purely quantitative or sampling-based validation. Evaluated on two benchmark datasets, it significantly improves downstream task performance while enabling efficient identification of systematic labeling errors and enhancing model generalization.

0 citationsRead paper

VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis

May 06, 2025

In safety-critical applications such as autonomous driving, validating object detection models via data slicing remains challenging due to reliance on metadata, insufficient concept-level annotations, high expert cognitive overhead, and lack of interactive hypothesis-testing mechanisms. Method: This paper introduces the first XAI framework integrating multimodal foundation models (CLIP and SAM) with visual analytics. It enables unsupervised visual slice discovery without requiring auxiliary image metadata or manual annotations; employs natural language generation (NLG) to produce interpretable, natural-language insights; and supports expert-driven, natural-language-based interactive hypothesis validation. Results: Evaluated across three real-world use cases and an expert study, the framework significantly reduces cognitive load, improves efficiency in identifying vulnerable slices, and comprehensively supports pre-deployment reliability assessment of object detection models.

0 citationsRead paper
Recent publications

Latest Papers

Observability for Delegated Execution in Agentic AI Systems

Jun 08, 2026

This work addresses the challenge that standard observability data in large language model agent systems often fails to distinguish execution traces across different delegation assignments, rendering cross-tool and cross-system delegation behaviors untraceable. To overcome this limitation, the paper proposes an agent-aware observability infrastructure that leverages a lightweight gateway and a unified information model to bind delegation context at runtime. This approach enables, for the first time, precise reconstruction of delegation scopes without relying on heuristic time windows. By supporting fine-grained behavioral forensics and direct forensic queries, the method significantly enhances observability and auditability of delegated executions in heterogeneous multi-agent systems.

0 citationsRead paper

Policy-Guided Threat Hunting: An LLM enabled Framework with Splunk SOC Triage

Mar 25, 2026

This work proposes an automated, dynamic threat-hunting framework that integrates AI agents with the Splunk SIEM platform to address the evolving challenges posed by advanced persistent threats (APTs) and the inefficiencies in processing massive volumes of heterogeneous logs within security operations centers (SOCs). The framework uniquely combines a reconstruction-based autoencoder, a two-layer deep reinforcement learning architecture, and a large language model to establish a policy-guided, context-aware autonomous hunting mechanism. This enables an end-to-end closed-loop pipeline—from traffic ingestion and anomaly detection to risk prioritization. Experimental results demonstrate that the approach adaptively aligns with diverse SOC objectives and effectively identifies suspicious and malicious network traffic on both public and simulated datasets, significantly enhancing analysts’ decision-making efficiency.

0 citationsRead paper

VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

Jul 11, 2025

Large-scale labels automatically generated by multimodal foundation models (e.g., CLIP, LLaVA) lack ground-truth annotations; existing evaluation methods rely on limited metrics or small-sample inspections, hindering detection of latent errors—especially in open-vocabulary image segmentation. Method: We propose the first human-in-the-loop visual analytics framework tailored for this task. It integrates multimodal output analysis, visual clustering-based diagnosis, interactive label correction, and an expert feedback loop to enable fine-grained quality assessment and iterative refinement. Contribution/Results: By introducing visual analytics into the quality assurance pipeline for auto-generated labels, our approach overcomes the limitations of purely quantitative or sampling-based validation. Evaluated on two benchmark datasets, it significantly improves downstream task performance while enabling efficient identification of systematic labeling errors and enhancing model generalization.

0 citationsRead paper

VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis

May 06, 2025

In safety-critical applications such as autonomous driving, validating object detection models via data slicing remains challenging due to reliance on metadata, insufficient concept-level annotations, high expert cognitive overhead, and lack of interactive hypothesis-testing mechanisms. Method: This paper introduces the first XAI framework integrating multimodal foundation models (CLIP and SAM) with visual analytics. It enables unsupervised visual slice discovery without requiring auxiliary image metadata or manual annotations; employs natural language generation (NLG) to produce interpretable, natural-language insights; and supports expert-driven, natural-language-based interactive hypothesis validation. Results: Evaluated across three real-world use cases and an expert study, the framework significantly reduces cognitive load, improves efficiency in identifying vulnerable slices, and comprehensively supports pre-deployment reliability assessment of object detection models.

0 citationsRead paper