| name | ai-redteam |
| description | Guides adversarial testing of AI systems—prompt injection, jailbreaks, tool abuse, data exfiltration,
bias and harmful output probes, multi-turn attacks, and automated red-team harnesses for LLM applications.
Use when red-teaming chatbots, agents, RAG systems, or copilots before launch, designing safety eval
suites, reproducing reported vulnerabilities, or validating mitigations after incidents—not for writing
corporate AI policy (ai-risk-governance), building production features (ai-engineer), or general
network or app penetration testing (penetration-tester, network-pentester), enterprise adversary
simulation or purple-team campaigns (red-team-specialist), authorized web/API OWASP testing
(web-pentester), or binary/firmware RE (reverse-engineer). Production safeguard
serving and gateways:
ml-infrastructure-engineer-safeguards. Safety classifier R&D and benchmarks:
ml-research-engineer-safeguards.
|
AI Red Team
When to Use
- Red-teaming chatbots, agents, RAG systems, or copilots before launch
- Designing safety evaluation suites and adversarial test harnesses
- Reproducing reported prompt injection or jailbreak vulnerabilities
- Validating mitigations after incidents (retesting filters, hardening)
- Running multi-turn coercion, encoding, or indirect injection campaigns
- Assessing bias, harmful output, or data exfiltration risks in LLM applications
- Scoping rules of engagement and severity rubrics for AI security testing
When NOT to Use
- Writing corporate AI policy or risk governance frameworks →
ai-risk-governance
- Building production LLM features or RAG pipelines →
ai-engineer
- General network/AD/infra penetration testing →
network-pentester
- Authorized web/API OWASP testing (non-LLM) →
web-pentester
- Enterprise adversary simulation, MITRE ATT&CK campaigns, purple team →
red-team-specialist
- Binary, firmware, or protocol reverse engineering →
reverse-engineer
- CI/CD pipeline security →
devsecops
Related skills
| Need | Skill |
|---|
| Production architecture and mitigations | ai-engineer |
| Governance sign-off and risk tiers | ai-risk-governance |
| Prompt design baselines | prompt-engineer |
| CI pipeline security | devsecops |
| Web/API OWASP pentest (non-LLM) | web-pentester |
| Network/AD/infra pentest (non-LLM) | network-pentester |
| Multi-domain pentest (non-LLM) | penetration-tester |
| Enterprise red team / adversary simulation (non-LLM) | red-team-specialist |
| Security program and pentest governance | cybersecurity |
| Deploy/monitor safeguard inference path | ml-infrastructure-engineer-safeguards |
| Safety benchmarks and classifier training | ml-research-engineer-safeguards |
| Post-incident disk/memory/log forensics and chain of custody | digital-forensics-analyst |
| Binary/protocol RE on non-LLM malware or implants | reverse-engineer |
| Security incident coordination after AI abuse | incident-responder |
Core Workflows
1. Scope and rules of engagement
- Define target: model, app surface, tools, data stores
- Obtain written authorization and time window
- Agree out-of-scope (e.g., no social engineering of employees unless approved)
- Define success criteria: critical findings, reproduction steps, severity rubric
- Plan safe test environment (no prod customer data)
See references/engagement_scope.md for ROE template and severity definitions.
2. Threat model for LLM applications
| Class | Examples |
|---|
| Prompt injection | Instructions in user/doc content override system policy |
| Jailbreak | Role-play, encoding, multi-turn coercion |
| Tool abuse | Unauthorized API calls, parameter injection |
| Data exfiltration | RAG leaks other tenants' chunks, PII in logs |
| Supply chain | Malicious tool definitions, compromised plugins |
| Denial of service | Token burn, recursive agent loops |
See references/attack_catalog.md for technique families and test prompts (use ethically).
3. Test execution
Phases:
- Baseline — document intended refusals and allowed behaviors
- Automated sweep — harness with curated attack set + fuzz mutations
- Manual creativity — domain-specific abuse scenarios
- Tool/RAG focus — indirect injection via retrieved documents
- Regression — re-run after mitigations
Log: input, output, tool calls, latency, whether guardrail fired.
See references/testing_harness.md for harness design and datasets.
4. Reporting
Each finding includes:
- Title and severity (impact × likelihood)
- Steps to reproduce (minimal)
- Evidence (redacted transcripts)
- Affected component
- Recommended mitigation
- Retest criteria
See references/reporting.md for report template and remediation tracking.
5. Mitigation validation
| Mitigation | Retest |
|---|
| Input/output filters | Bypass attempts with paraphrases |
| System prompt hardening | Injection via RAG context |
| Tool allowlists | Confused deputy and scope creep |
| Human approval gate | Automated agent bypass paths |
See references/mitigations.md for defense depth and known weak controls.
When to load references
- ROE and scope →
references/engagement_scope.md
- Attack types →
references/attack_catalog.md
- Harness and automation →
references/testing_harness.md
- Reports →
references/reporting.md
- Defenses →
references/mitigations.md