| name | ai-gateway-guardrails |
| description | Enforce Input/Output Guardrails at the LLM Gateway layer โ PII redaction, Prompt Injection defense, Jailbreak detection, Toxicity filter, and Tool Allow-list. Integrates Bedrock Guardrails, NeMo Guardrails, Llama Guard 3, and regex/regex-ML policies on Bifrost/LiteLLM with Langfuse audit trail. |
| argument-hint | [compliance scope โ ISMS-P, finance, healthcare] |
| user-invocable | true |
| model | claude-sonnet-4-6 |
| allowed-tools | Read,Write,Edit,Bash,Grep,Glob,mcp__eks,mcp__aws-documentation,mcp__well-architected-security |
When to Use
- ํ๊ตญ ๊ธ์ต๊ถ(์ ์๊ธ์ต๊ฐ๋
๊ท์ ยทISMS-P), ์๋ฃ, ๊ณต๊ณต ๋ฑ ๊ท์ ํ๊ฒฝ์ LLM ์๋น์ค๋ฅผ ๋ฐฐํฌํ ๋
- Prompt Injection / Jailbreak / PII ์ ์ถ / Tool Poisoning ์ํ์ ๋ฐฉ์ดํด์ผ ํ ๋
- Bedrock Guardrails, NeMo Guardrails, Llama Guard 3 ์ค ์ ํ ๋ฐ ์กฐํฉ์ด ํ์ํ ๋
- Agent ๊ฐ ์ธ๋ถ Tool ์ ํธ์ถํ ๋ Allow-list ๊ธฐ๋ฐ ์ ์ฑ
์ด ํ์ํ ๋
When NOT to Use
- ๋ด๋ถ PoC ๋ก ์ํ ๋ชจ๋ธ์ด ๋ถํ์ โ Guardrail ์ค๋ฒํค๋๋ง ๋ฐ์
- Bedrock ๋งค๋์ง๋ ๋ชจ๋ธ๋ง ํธ์ถํ๋ฉฐ Bedrock Guardrails ๊ธฐ๋ณธ ํ์ฑ โ ์ถ๊ฐ ๊ตฌ์ฑ ๋ถํ์ (๋จ, ๋ก๊ทธ๋ ํ์)
- ๋ด๋ถ RAG ์์ด ๋จ์ Q&A โ regex ์์ค์ Input Guard ๋ง ํ์ํ ์ ์์
Preconditions
- Inference Gateway (Bifrost/LiteLLM) ๊ฐ ์ด๋ฏธ ๋ฐฐํฌ๋จ (
inference-gateway-routing ์๋ฃ)
- Langfuse ๊ฐ audit log ๋ฅผ ์์ ๊ฐ๋ฅ (
langfuse-observability ์๋ฃ)
- PII ์ ์ฑ
ยท์ฐจ๋จ ์นดํ
๊ณ ๋ฆฌยทTool Allow-list ์ ์ ๋ฌธ์ ํ๋ณด
Procedure
Step 1. ์ํ ๋ชจ๋ธ ์ ์ (OWASP LLM Top 10 ๊ธฐ๋ฐ)
- LLM01 Prompt Injection (Direct/Indirect)
- LLM02 Sensitive Information Disclosure (PII, ์์
๋น๋ฐ)
- LLM06 Excessive Agency (Tool ์ค์ฉ)
- LLM08 Vector & Embedding Weaknesses (RAG poisoning)
Step 2. ๋ค์ธต ๋ฐฉ์ด (Defense in Depth)
User โ Input Guard โ Gateway Policy โ Tool Allow-list โ LLM โ Output Guard โ Response
โ
Audit Log (Langfuse)
- Input Guard: PII redaction, Injection pattern, Jailbreak classifier
- Gateway Policy: AuthN/Z, Rate Limit, Tenant Isolation
- Tool Allow-list: MCP Server Registry, Scoped tokens
- Output Guard: PII scrub, Toxicity, Fact check
- Audit Log: ๋ชจ๋ ๋จ๊ณ์์ Langfuse + CloudTrail ๊ธฐ๋ก
Step 3. Bedrock Guardrails ์ฐ๋ (๋งค๋์ง๋)
providers:
bedrock:
region: ap-northeast-2
guardrails:
- id: arn:aws:bedrock:ap-northeast-2:ACCOUNT:guardrail/PII-BLOCK
version: "1"
- id: arn:aws:bedrock:ap-northeast-2:ACCOUNT:guardrail/TOXICITY
version: "1"
Step 4. NeMo Guardrails (์คํ์์ค Flow)
models:
- type: main
engine: openai
model: gpt-4.1
rails:
input:
flows:
- self check input
- detect pii
output:
flows:
- self check output
- remove pii
- fact checking
Step 5. Llama Guard 3 (Output Classifier)
- Meta Llama Guard 3 8B ๋ชจ๋ธ์ vLLM ๋ณ๋ Pod ๋ก ๋ฐฐํฌ
- Bifrost output ํ
์์ Llama Guard 3 call โ unsafe ํ์ ์ ์ฌ์์ฑ ๋๋ ์ฐจ๋จ
Step 6. Tool Allow-list (MCP)
mcpAllowList:
- name: aws-documentation
scopes: ["read"]
- name: eks
scopes: ["read", "describe"]
tokenPolicy:
maxLifetimeSeconds: 900
audience: ai-infra
Step 7. Audit & ์๋ฆผ
- ๋ชจ๋ guard violation ์ Langfuse
scores + tags ๋ก ๊ธฐ๋ก
- Prometheus ๋ฉํธ๋ฆญ
guardrail_violation_total{type="pii",decision="block"}
- CloudWatch Logs + SIEM ์ฐ๊ณ (Security Lake)
- Slack/PagerDuty ์๋ฆผ ๊ธฐ์ค:
guardrail_violation_rate > 5%/5m
Good Examples
- ISMS-P ๋์ ๊ธ์ต: Bedrock Guardrails(managed PII + Block) + NeMo Guardrails(์์ฒด Policy) + Llama Guard 3(output)
- Coding Agent: Tool Allow-list ๋ก
shell_exec, network_request ์ฐจ๋จ
- RAG: Indirect Injection ๋ฐฉ์ด์ฉ Llama Guard 3 + fact-check Flow
Bad Examples (๊ธ์ง)
- Guardrails ์์ด Tool-calling Agent ๋ฅผ ํ๋ก๋์
๋ฐฐํฌ โ LLM06 ์ฆ์ ์๋ฐ
- ์ ๊ท์ ๊ธฐ๋ฐ PII ๋จ๋
โ ํ๊ตญ ์ฃผ๋ฏผ๋ฒํธ ๋ณํ ํจํด ๋ฏธํ์ง, ML classifier ๋ณํ ํ์
- Audit log ๋ฏธ์์ง โ ๊ท์ ๊ฐ์ฌ ์ ๊ทผ๊ฑฐ ๋ถ์ฌ
allowed-tools: ["*"] โ ์ ์ฒด ํ์ฉ = ์ ์ฑ
์์
References