Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM Guard's PromptInjection scanner or Hugging Face Prompt Guard 2. Use when an agent ingests untrusted external content and you need to screen it for injected instructions before the LLM processes it.
Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.
Une commande directe contourne le prompt de vérification. Examinez la source avant de l'exécuter.
Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM Guard's PromptInjection scanner or Hugging Face Prompt Guard 2. Use when an agent ingests untrusted external content and you need to screen it for injected instructions before the LLM processes it.
Authorized-use-only notice: Scripts in this skill scan untrusted content for injection payloads and run detector models. Run scanning only on data you are authorized to process, and treat any extracted payloads as live untrusted input — never paste them back into a privileged LLM context.
Overview
Indirect prompt injection (MITRE ATLAS AML.T0051.001, OWASP LLM01:2025) occurs when an LLM-powered agent ingests external content — a web page it browses, a PDF or email it summarizes, an image it OCRs, a tool result it reads — and that content contains hidden instructions the model then follows as if they came from the developer or user. Because the agent treats all tokens in its context window as equally authoritative, an attacker who controls any consumed artifact can hijack the agent's behavior: exfiltrate conversation history, redirect tool calls, leak secrets, or pivot through connected systems.
Unlike direct injection (the user types the attack), indirect injection arrives through a trusted-looking data channel, which is why naive input filtering misses it. Payloads hide in many forms: HTML comments and display:none/zero-width text on web pages, white-on-white or tiny-font text in PDFs, alt-text and EXIF metadata in images, text rendered into pixels (invisible to OCR-light filters but read by multimodal models), Unicode tag/zero-width characters, and Base64/ROT13 obfuscation. This skill builds a detection pipeline that normalizes and scans every artifact before it reaches the model, combining heuristic/regex detection, dedicated detector models (Meta Prompt Guard 2, ProtectAI's deberta-v3 prompt-injection classifier via LLM Guard), and multimodal extraction for images, and then defines response actions and detection telemetry.
When to Use
When building or hardening an agent that browses the web, reads email, summarizes documents, or processes user-uploaded files/images.
When you need a content-sanitization gate in front of an LLM that ingests third-party data.
During AI red-team / blue-team exercises validating that injected instructions in retrieved artifacts are caught.
When investigating an incident where an agent behaved as if it received instructions you did not author.
As a CI/CD pre-ingestion scan for documents added to a knowledge base.
Run heuristic and ML-based injection detectors (LLM Guard PromptInjection scanner, Prompt Guard 2).
Score each artifact and enforce a block / sanitize / allow decision before model ingestion.
Emit structured detection telemetry suitable for a SIEM and map findings to ATLAS AML.T0051.001.
MITRE ATT&CK Mapping
ID
Official Name
Relevance
AML.T0051.001
LLM Prompt Injection: Indirect
The exact technique this skill detects and mitigates
AML.T0051
LLM Prompt Injection
Parent technique covering all prompt-injection variants
AML.T0057
LLM Data Leakage
Common objective of an indirect injection that this detection prevents
AML.T0053
LLM Plugin Compromise
Injected instructions frequently target the agent's tools/plugins
Workflow
1. Extract hidden text from web content
Pull comments, hidden elements, and metadata that a human never sees but the model does.
# extract_html.pyfrom bs4 import BeautifulSoup, Comment
defextract_hidden(html: str):
soup = BeautifulSoup(html, "html.parser")
hidden = []
for c in soup.find_all(string=lambda t: isinstance(t, Comment)):
hidden.append(("comment", c.strip()))
for el in soup.select('[style*="display:none"],[style*="visibility:hidden"],[hidden]'):
hidden.append(("css-hidden", el.get_text(strip=True)))
for img in soup.find_all("img"):
if img.get("alt"):
hidden.append(("alt-text", img["alt"]))
return [h for h in hidden if h[1]]
2. Normalize and de-obfuscate
Strip zero-width / Unicode-tag characters and decode common encodings so detectors see the real payload.
# normalize.pyimport base64, codecs, re, unicodedata
ZERO_WIDTH = dict.fromkeys(map(ord, ""), None)
TAG_RANGE = range(0xE0000, 0xE0080) # Unicode tag chars used to smuggle textdefnormalize(text: str) -> str:
text = text.translate(ZERO_WIDTH)
text = "".join(ch for ch in text iford(ch) notin TAG_RANGE)
text = unicodedata.normalize("NFKC", text)
for token in re.findall(r"[A-Za-z0-9+/=]{20,}", text):
try:
decoded = base64.b64decode(token).decode("utf-8", "ignore")
if decoded.isprintable():
text += f"\n[decoded-b64] {decoded}"except Exception:
pass
text += "\n[decoded-rot13] " + codecs.decode(text, "rot_13")
return text
3. Scan with LLM Guard's PromptInjection scanner
LLM Guard wraps a transformer classifier and returns a risk score per input.
Run the pipeline over a labeled set of clean + injected artifacts, measure precision/recall, and tune threshold to balance false positives against missed injections. Re-test whenever the agent's model or ingestion sources change.