| name | detecting-ai-model-prompt-injection-attacks |
| description | Detects prompt injection attacks targeting LLM-based applications using a multi-layered defense combining regex pattern matching for known attack signatures, heuristic scoring for structural anomalies, and transformer-based classification with DeBERTa models. The detector analyzes user inputs before they reach the LLM, flagging direct injections (system prompt overrides, role-play escapes, instruction hijacking) and indirect injections (encoded payloads, multi-language obfuscation, delimiter-based escapes). Based on the OWASP LLM Top 10 (LLM01:2025 Prompt Injection) and Simon Willison's prompt injection taxonomy. Activates for requests involving prompt injection detection, LLM input sanitization, AI security scanning, or prompt attack classification.
|
| domain | cybersecurity |
| subdomain | ai-security |
| tags | ["prompt-injection","LLM-security","OWASP-LLM-Top10","NLP-classification","input-validation"] |
| version | 1.0.0 |
| author | mukul975 |
| license | Apache-2.0 |
| atlas_techniques | ["AML.T0051","AML.T0054","AML.T0056","AML.T0068","AML.T0067"] |
| nist_ai_rmf | ["GOVERN-1.1","GOVERN-6.1","MEASURE-2.7","MEASURE-2.5","MANAGE-2.4"] |
| d3fend_techniques | ["Content Validation","Content Filtering","Application Hardening","Inbound Traffic Filtering","User Behavior Analysis"] |
| nist_csf | ["GV.OC-03","ID.RA-01","PR.PS-01","DE.AE-02"] |
Detecting AI Model Prompt Injection Attacks
When to Use
- Scanning user inputs to LLM-powered applications before they are forwarded to the model
- Building an input validation layer for chatbots, AI agents, or retrieval-augmented generation (RAG) pipelines
- Monitoring logs of LLM interactions to retrospectively identify prompt injection attempts
- Evaluating the effectiveness of existing prompt injection defenses through red-team testing
- Classifying prompt injection payloads during security incident investigations involving AI systems
Do not use as the sole defense mechanism against prompt injection -- always combine with output validation, privilege separation, and least-privilege tool access. Not suitable for detecting jailbreaks that do not involve injection of adversarial instructions.
Detection Gaps & Validation
Prompt-injection detectors built on regex + classifier most often miss attacks
that never look like the canonical "ignore previous instructions":
- Obfuscated / encoded payloads: base64, ROT13, hex, leetspeak, zero-width
characters, or homoglyphs carry the instruction past signature regexes.
Decode-then-rescan, and test with
"aWdub3JlIGFsbCBydWxlcw==" style inputs.
- Indirect / cross-context injection: the malicious instruction arrives via
RAG-retrieved documents, tool/API output, or webpage content the model
ingests - not the user field your filter watches. Validate by planting an
injected instruction inside a retrieved document and confirming the detector
sees it.
- Multilingual evasion: an instruction in a low-resource language, or mixed
script, slips an English-trained classifier. Test non-English jailbreaks.
- Payload splitting / accretion: the attack is assembled across turns or
concatenated fragments, each benign alone. Test multi-turn assembly.
- How to validate detection fires + tune FPs: run a labeled red-team corpus
(deepset/prompt-injections plus encoded/indirect/multilingual variants),
confirm true positives trip at the configured threshold, and measure false
positives against benign code snippets and technical text - lower the
threshold or add layers until both error rates are acceptable.
Prerequisites
- Python 3.10+ with pip for installing detection dependencies
- The
transformers and torch libraries for running the DeBERTa-based classifier model
- The
protectai/deberta-v3-base-prompt-injection-v2 model from Hugging Face (downloaded on first run, approximately 700 MB)
- Network access to Hugging Face Hub for initial model download (offline mode supported after first download)