Detecting Prompt Injection Attacks Against Application Using Classifiers

📅 2025-12-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address prompt injection attacks threatening large language model (LLM)-powered web applications, this paper systematically constructs the first high-quality, real-world-oriented prompt injection benchmark dataset—enhanced from the HackAPrompt Playground—and empirically evaluates four classifier families: LSTM, feedforward neural networks, random forests, and naive Bayes. We propose a lightweight detection framework integrating data augmentation and binary classification, achieving significant improvements in detection accuracy and cross-scenario generalization while maintaining low inference overhead. Our key contributions are: (1) the first publicly released, high-fidelity prompt injection benchmark dataset; (2) an empirical characterization of performance boundaries of diverse detectors under realistic deployment conditions; and (3) a plug-and-play, low-latency security module that enables real-time defense—providing actionable, production-ready safeguards for securing LLM-integrated applications.

Technology Category

Application Category

📝 Abstract
Prompt injection attacks can compromise the security and stability of critical systems, from infrastructure to large web applications. This work curates and augments a prompt injection dataset based on the HackAPrompt Playground Submissions corpus and trains several classifiers, including LSTM, feed forward neural networks, Random Forest, and Naive Bayes, to detect malicious prompts in LLM integrated web applications. The proposed approach improves prompt injection detection and mitigation, helping protect targeted applications and systems.
Problem

Research questions and friction points this paper is trying to address.

Detects prompt injection attacks in LLM-integrated applications
Trains classifiers to identify malicious prompts in web systems
Improves security by enhancing prompt injection detection methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Augmented dataset from HackAPrompt corpus
Trained multiple classifiers including LSTM and Random Forest
Detected malicious prompts in LLM web applications
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Safwan Shaheer
Department of Computer Science (CS), School of Data and Sciences (SDS), Dhaka, Bangladesh
G
G. M. Refatul Islam
Department of Computer Science and Engineering (CSE), School of Data and Sciences (SDS), Dhaka, Bangladesh
M
Mohammad Rafid Hamid
Department of Computer Science and Engineering (CSE), School of Data and Sciences (SDS), Dhaka, Bangladesh
M
Md. Abrar Faiaz Khan
Department of Computer Science and Engineering (CSE), School of Data and Sciences (SDS), Dhaka, Bangladesh
M
Md. Omar Faruk
Department of Computer Science and Engineering (CSE), School of Data and Sciences (SDS), Dhaka, Bangladesh
Y
Yaseen Nur
Department of Computer Science (CS), School of Data and Sciences (SDS), Dhaka, Bangladesh