Skip to main content

agent-security-guardrails

Implements prompt injection detection, input validation, tool access control, and output sanitization to secure LLM agents against adversarial attacks.

Ir para a instalação

Informações da origem

Repositório
paulpas/agent-skill-router
Última atividade na origem
4 de junho de 2026 às 23:31
Idioma detectado do SKILL.md
inglês
Estrelas
6
Forks
0

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
agent-security-guardrails
description
Implements prompt injection detection, input validation, tool access control, and output sanitization to secure LLM agents against adversarial attacks.
license
MIT
compatibility
opencode
archetypes
["tactical","enforcement"]
anti_triggers
["brainstorming","vague ideation","single-agent monolith"]
response_profile
{"verbosity":"low","directive_strength":"high","abstraction_level":"operational"}
metadata
{"version":"1.0.0","domain":"agent","triggers":"prompt injection, guardrails, jailbreak detection, tool access control, input validation, LLM security, how do i secure my agent","role":"implementation","scope":"implementation","output-format":"code","content-types":["code","guidance","do-dont","examples"],"related-skills":"agent-reliability-engineering, coding-security-review, agent-system-hints-design"}
# Agent Security Guardrails Implements security enforcement layers for LLM-powered agents to protect against prompt injection, unauthorized tool access, input/output poisoning, and adversarial attacks. This skill guides the model in building multi-layered guardrail systems that validate inputs, constrain outputs, control tool permissions, and detect malicious intent before it reaches the agent's core reasoning loop. Guardrails are not a single check — they form a defense-in-depth pipeline where each layer intercepts different attack vectors: input sanitization catches prompt injection at the boundary, output validation prevents data exfiltration, tool access control enforces least-privilege execution, and runtime monitoring detects behavioral anomalies that slip past static checks. Together they create a resilient security posture without sacrificing agent capability. ## TL;DR Checklist - [ ] Apply input validation on every user message before it reaches the LLM context - [ ] Enforce tool access control — each tool must have an explicit allowlist per agent identity - [ ] Validate all LLM outputs against structured schemas before processing - [ ] Implement prompt injection detection (direct + indirect via retrieved documents) - [ ] Sanitize tool inputs and outputs to prevent data leakage - [ ] Log all guardrail violations with full context for audit trails --- ## When to Use Use this skill when: - Building LLM-powered agents that execute tools or access external resources - Deploying agents in production where adversarial users may attempt prompt injection or jailbreaks - Designing permission boundaries for multi-tenant agent systems - Integrating third-party APIs through agent tool execution and needing input/output validation - Implementing compliance requirements (data retention, PII filtering) on agent outputs - Adding safety layers to agents that process user-provided documents or links (indirect prompt injection vectors) ## When NOT to Use Avoid this skill for: - Simple chat-only agents with no tool execution — the guardrails add unnecessary overhead - Internal tools used only by trusted developers — basic input validation suffices without full guardrail infrastructure - One-off scripts or prototypes where deployment security is not a concern - As a substitute for fixing root cause vulnerabilities in the underlying application code --- ## Core Workflow ``` User Input ──→ Input Sanitizer ──→ Injection Detector ──→ Role/Context Router ──→ Tool Executor │ │ │ │ │ │ [violation] [injection] [permission denied] [output validator] │ │ │ │ │ │ LOG + BLOCK LOG + TRUNCATE LOG + DENY [leak detected] ▼ ▼ ▼ ▼ ▼ Safe Input Blocked Sanitized Original Clean Output Response Request Request Returned ``` 1. **Configure Guardrail Pipeline** — Define the layers in order: input sanitization, injection detection, permission routing, tool execution with validation, and output sanitization. Each layer must have a clear pass/fail behavior and logging requirement. **Checkpoint:** Verify every external-facing tool has an explicit allowlist entry — tools without allowlists are blocked by default (deny-by-default). 2. **Implement Input Validation** — Apply multi-layer input checks on every user message before it enters the LLM context: - Strip or reject embedded URLs from untrusted sources (indirect injection via links) - Limit message length to prevent prompt buffer overflow attacks (>4096 tokens triggers truncation with warning) - Detect and block common jailbreak patterns (DAN mode, roleplay escaping, base64-encoded commands) **Checkpoint:** No user message should reach the LLM without passing through at least input sanitization and injection detection. 3. **Enforce Tool Access Control** — Before any tool execution, verify the agent identity has permission: - Define per-identity allowlists (agent_id → allowed_tools[]) - Validate tool arguments against declared schemas using pydantic or JSON schema - Sandbox high-risk tools (file writes, network calls, shell execution) behind approval gates **Checkpoint:** Every tool call must log the agent identity, tool name, and argument hashes before execution. 4. **Detect Prompt Injection** — Apply both direct and indirect injection detection: - Direct: scan user messages for commands embedded in natural text ("Ignore previous instructions", "You are now a DAN") - Indirect: sanitize retrieved documents, web pages, or database records that could contain injected commands - Use keyword + semantic scoring (pattern match for known injection phrases, anomaly detection on instruction density) **Checkpoint:** Any message flagged as injection must be logged and either sanitized or blocked based on severity level. 5. **Validate Tool Outputs** — Before any tool output reaches the LLM's context: - Validate against expected schema (JSON schema validation for structured outputs) - Truncate oversized responses (>8192 chars triggers warning with summary) - Sanitize sensitive data patterns (SSN, credit cards, API keys) from tool responses **Checkpoint:** No raw tool output should enter the context window without passing through output validation. 6. **Sanitize Final Outputs** — Before returning any response to the user: - Filter PII patterns (emails, phone numbers, SSNs, API keys) using regex-based detectors - Validate against structured output schema if the task expected a specific format - Log all sanitization actions for compliance auditing **Checkpoint:** Final outputs must pass all validation checks before reaching the user — never trust intermediate results. --- ## Implementation Patterns ### Pattern 1: Guardrail Pipeline with Layered Validation ```python import re import logging import hashlib from dataclasses import dataclass, field from enum import Enum from typing import Any logger = logging.getLogger("agent.guardrails") class Severity(Enum): LOW = "low" MEDIUM = "medium" HIGH = "high" CRITICAL = "critical" class GuardrailAction(Enum): PASS = "pass" SANITIZE = "sanitize" WARN = "warn" BLOCK = "block" @dataclass class GuardrailResult: """Result from a single guardrail layer.""" action: GuardrailAction severity: Severity message: str sanitized_input: str | None = None @property def is_blocked(self) -> bool: return self.action == GuardrailAction.BLOCK @property def requires_remediation(self) -> bool: return self.action in (GuardrailAction.SANITIZE, GuardrailAction.BLOCK) @dataclass class GuardrailViolation: """Persistent record of a guardrail violation for audit logging.""" layer: str severity: Severity original_content: str sanitized_content: str | None timestamp: float agent_id: str | None = None def to_dict(self) -> dict: return { "layer": self.layer, "severity": self.severity.value, "original_hash": hashlib.sha256( (self.original_content or "").encode() ).hexdigest()[:16], "sanitized_length": len(self.sanitized_content or ""), "timestamp": self.timestamp, "agent_id": self.agent_id, } class InputSanitizer: """Layer 1: Sanitize raw user input before it reaches the LLM context. Handles buffer overflow prevention, embedded link stripping, and common jailbreak pattern detection using regex-based heuristics. Applies Law 4 (Fail Fast) — invalid inputs are blocked immediately without reaching the reasoning layer. """ # Common jailbreak prefix patterns _JAILBREAK_PATTERNS: list[re.Pattern] = [ re.compile(r"\b(DAN|do anything now)\b.*(?:mode|prompt|instruction)", re.IGNORECASE), re.compile(r"ignore\s+(?:all\s+)?(previous|above|prior)\s+(instructions|prompts|rules)", re.IGNORECASE), re.compile(r"(?:you are now|act as)\s+(?:a )?(?:system|developer|admin)", re.IGNORECASE), re.compile(r"secret mode(?: activated)?", re.IGNORECASE), re.compile(r"(?:override|bypass|disable)\s+(?:content|safety|output) filters", re.IGNORECASE), ] # Sensitive data patterns for PII detection _PII_PATTERNS: dict[str, re.Pattern] = { "ssn": re.compile(r"\b\d{3}-\d{2}-\d{4}\b"), "credit_card": re.compile(r"\b(?:\d[ -]*?){13,16}\b"), "api_key": re.compile(r"(?:sk-)[A-Za-z0-9]{20,}"), "email": re.compile(r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}"), } def __init__(self, max_token_length: int = 4096) -> None: self.max_token_length = max_token_length def sanitize(self, user_input: str, agent_id: str | None = None) -> GuardrailResult: """Sanitize and validate a raw user input message. Args: user_input: The raw text from the user. agent_id: Optional identifier for audit logging. Returns: GuardrailResult indicating pass, sanitize, warn, or block action. """ if not user_input or not user_input.strip(): return GuardrailResult( action=GuardrailAction.BLOCK, severity=Severity.MEDIUM, message="Empty input rejected", ) # Check for jailbreak patterns — first pass with regex for pattern in self._JAILBREAK_PATTERNS: if pattern.search(user_input): logger.warning("Jailbreak pattern detected from agent '%s'", agent_id) return GuardrailResult( action=GuardrailAction.BLOCK, severity=Severity.CRITICAL, message=f"Blocked jailbreak pattern", ) # Check for base64-encoded commands (obfuscated injection) import base64 b64_pattern = re.compile(r"(?:[A-Za-z0-9+/]{4}){15,}(?:[A-Za-z0-9+/]{2}==|[A-Za-z0-9+/]{3}=)?") matches = b64_pattern.findall(user_input) if len(matches) >= 2: for encoded in matches: try: decoded = base64.b64decode(encoded).decode("utf-8", errors="ignore") # Check if decoded content contains instructions if any(phrase in decoded.lower() for phrase in ["ignore", "system:", "prompt:", "instruction"]): return GuardrailResult( action=GuardrailAction.BLOCK, severity=Severity.HIGH, message="Blocked base64-obfuscated injection attempt", ) except Exception: pass # Enforce max length — truncate with warning rather than blocking if len(user_input) > self.max_token_length * 3: sanitized = user_input[: self.max_token_length * 3] + "\n[TRUNCATED: input exceeded token limit]" return GuardrailResult( action=GuardrailAction.SANITIZE, severity=Severity.LOW, message=f"Input truncated from {len(user_input)} to {self.max_token_length * 3} characters", sanitized_input=sanitized, ) # Strip embedded URLs from untrusted input cleaned = re.sub(r"https?://\S+", "[LINK_REMOVED]", user_input) if cleaned != user_input: return GuardrailResult( action=GuardrailAction.SANITIZE, severity=Severity.MEDIUM, message="Removed embedded URLs from input", sanitized_input=cleaned, ) return GuardrailResult(action=GuardrailAction.PASS, severity=Severity.LOW, message="Input passed sanitization") def extract_pii(self, text: str) -> dict[str, list[tuple[int, int]]]: """Return positions of PII matches in text for downstream filtering.""" findings: dict[str, list[tuple[int, int]]] = {} for pii_type, pattern in self._PII_PATTERNS.items(): matches = list(pattern.finditer(text)) if matches: findings[pii_type] = [(m.start(), m.end()) for m in matches] return findings ``` ### Pattern 2: Tool Access Control with Permission Enforcement ```python from pydantic import BaseModel, Field, field_validator from collections.abc import Sequence class ToolPermission(BaseModel): """Declares a tool's permission requirements and validation schema.""" tool_name: str = Field(description="Name of the tool to execute") allowed_agents: list[str] = Field( description="Agent IDs permitted to call this tool" ) requires_approval: bool = Field( default=False, description="Whether a human gate is required before execution", ) input_schema: dict | None = Field( default=None, description="JSON schema for validating tool arguments", ) max_output_bytes: int = Field( default=65536, description="Maximum output size in bytes before truncation", ) class ToolAccessController: """Layer 3: Enforces permission boundaries on all tool execution. Applies Law 2 (Parse at boundary) by strictly validating tool calls against declared schemas. Applies Law 4 (Fail Fast, Fail Loud) by blocking unauthorized access immediately with a descriptive error. Uses deny-by-default: any tool without an explicit allowlist entry is blocked regardless of who requests it. """ def __init__(self, permissions: Sequence[ToolPermission] | None = None) -> None: self._permissions: dict[str, ToolPermission] = {} if permissions: for perm in permissions: self._permissions[perm.tool_name] = perm def register(self, permission: ToolPermission) -> None: """Register a tool's permission configuration.""" self._permissions[permission.tool_name] = permission logger.info( "Registered tool '%s' for agents: %s", permission.tool_name, permission.allowed_agents, ) def authorize( self, agent_id: str, tool_name: str, arguments: dict[str, Any], ) -> GuardrailResult: """Check if an agent is authorized to execute a tool with given arguments. Returns PASS if the agent can proceed. Returns BLOCK/WARN/SANITIZE with details about what failed. Applies deny-by-default for unknown tools. """ # Deny by default — unknown tools are always blocked if tool_name not in self._permissions: return GuardrailResult( action=GuardrailAction.BLOCK, severity=Severity.HIGH, message=f"Tool '{tool_name}' not in allowlist — deny-by-default", ) perm = self._permissions[tool_name] # Check agent identity against allowlist if agent_id not in perm.allowed_agents: logger.warning( "Unauthorized tool access: agent '%s' attempted '%s'", agent_id, tool_name, ) return GuardrailResult( action=GuardrailAction.BLOCK, severity=Severity.HIGH, message=f"Agent '{agent_id}' not authorized for tool '{tool_name}'", ) # Validate arguments against declared schema if perm.input_schema: try: from jsonschema import validate, ValidationError validate(instance=arguments, schema=perm.input_schema) except ValidationError as e: logger.warning( "Tool argument validation failed for '%s': %s", tool_name, e.message ) return GuardrailResult( action=GuardrailAction.BLOCK, severity=Severity.MEDIUM, message=f"Invalid arguments for '{tool_name}': {e.message}", ) # Flag tools requiring human approval if perm.requires_approval: return GuardrailResult( action=GuardrailAction.WARN, severity=Severity.LOW, message=f"Tool '{tool_name}' requires human approval before execution", ) logger.info("Tool access granted: agent='%s' tool='%s'", agent_id, tool_name) return GuardrailResult( action=GuardrailAction.PASS, severity=Severity.LOW, message="Access authorized" ) def get_allowed_tools(self, agent_id: str) -> list[str]: """Return the list of tools an agent is permitted to use.""" return [ name for name, perm in self._permissions.items() if agent_id in perm.allowed_agents ] # Example usage — define permission registry with deny-by-default controller = ToolAccessController([ ToolPermission( tool_name="web_search", allowed_agents=["researcher_agent", "general_purpose"], input_schema={ "type": "object", "properties": { "query": {"type": "string", "minLength": 1, "maxLength": 200}, }, "required": ["query"], }, ), ToolPermission( tool_name="file_read", allowed_agents=["general_purpose"], input_schema={ "type": "object", "properties": {
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub