| name | detecting-ai-model-prompt-injection-attacks |
| description | Use when detects prompt injection attacks targeting LLM-based applications using a multi-layered defense combining regex pattern matching for known attack signatures, heuristic scoring for structural anomalies, and transformer-based classification with DeBERTa models. The detector analyzes user inputs before they reach the LLM, flagging direct injections (system prompt overrides, role-play escapes, instruction hijacking) and indirect injections (encoded payloads, multi-language obfuscation, d... |
| domain | cybersecurity |
| tags | ["prompt-injection","LLM-security","OWASP-LLM-Top10","NLP-classification","input-validation"] |
| subdomain | ai-security |
| version | 1.0.0 |
| author | oyi77 |
| license | Apache-2.0 |
| atlas_techniques | ["AML.T0051","AML.T0054","AML.T0056","AML.T0068","AML.T0067"] |
| nist_ai_rmf | ["GOVERN-1.1","GOVERN-6.1","MEASURE-2.7","MEASURE-2.5","MANAGE-2.4"] |
| d3fend_techniques | ["Content Validation","Content Filtering","Application Hardening","Inbound Traffic Filtering","User Behavior Analysis"] |
| nist_csf | ["GV.OC-03","ID.RA-01","PR.PS-01","DE.AE-02"] |
Detecting Ai Model Prompt Injection Attacks
Overview
Cybersecurity skill for detecting ai model prompt injection attacks. Follows industry best practices and security standards.
When to Use
Trigger phrases:
-
"detecting ai model prompt injection attacks"
-
"Detects prompt injection attacks targeting LLM-based applications using a multi-"
-
Scanning user inputs to LLM-powered applications before they are forwarded to the model
-
Building an input validation layer for chatbots, AI agents, or retrieval-augmented generation (RAG) pipelines
-
Monitoring logs of LLM interactions to retrospectively identify prompt injection attempts
-
Evaluating the effectiveness of existing prompt injection defenses through red-team testing
-
Classifying prompt injection payloads during security incident investigations involving AI systems
Do not use as the sole defense mechanism against prompt injection -- always combine with output validation, privilege separation, and least-privilege tool access. Not suitable for detecting jailbreaks that do not involve injection of adversarial instructions.
When NOT to Use
- When you lack proper authorization for testing
- For production systems without change management
- When the task requires legal or compliance expertise beyond technical scope
Prerequisites
- Python 3.10+ with pip for installing detection dependencies
- The
transformers and torch libraries for running the DeBERTa-based classifier model
- The protectai > deberta-v3-base-prompt-injection-v2 model from Hugging Face (downloaded on first run, approximately 700 MB)
- Network access to Hugging Face Hub for initial model download (offline mode supported after first download)
- Sample prompt injection payloads for testing (the script includes a built-in test suite)
Workflow
import re
IOC_PATTERNS = {
"ip": r"\b(?:\d{1,3}\.){3}\d{1,3}\b",
"domain": r"\b[a-z0-9-]+\.[a-z]{2,}\b",
"hash_md5": r"\b[a-f0-9]{32}\b",
"hash_sha256": ,
}
() -> :
{k: re.findall(v, text) k, v IOC_PATTERNS.items()}