- name
- foundry-security-spec
- description
- Implement Cisco's Foundry specification for agentic AI security evaluation systems with multi-agent architecture
- triggers
- ["implement foundry security evaluation","set up agentic security testing","create foundry spec system","build security evaluation agents","configure foundry security roles","design ai vulnerability discovery","implement foundry detector agents","create security finding workflow"]
# Foundry Security Spec
> Skill by [ara.so](https://ara.so) — Security Skills collection.
Foundry is an open specification from Cisco for building agentic AI security evaluation systems. It defines a multi-agent architecture with 8 core roles and 5 extension roles that coordinate to discover, validate, and report security findings. This is NOT a tool to install—it's a blueprint for building your own security evaluation system.
## Core Concepts
Foundry provides:
- **Architecture**: 8 core agent roles (Orchestrator, Planner, Navigator, Detector, Explorer, Validator, Investigator, Publisher)
- **Finding Lifecycle**: States, verdicts, evidence gates, fingerprinting
- **Coordination Model**: Atomic claims, heartbeat liveness, auto-blocking
- **Governance**: Sandboxing, budgets, yield-gated auto-stop, coverage gates
- **Detection-to-Prevention Flywheel**: Rules catch known issues, explorers find new ones, gaps become new rules
Works with CodeGuard rule format for portable detection rules that transfer between evaluation and prevention.
## Repository Structure
```
foundry-security-spec/
├── spec.md # Main specification (~130 functional requirements)
├── constitution.md # 11 inviolable principles
├── GLOSSARY.md # Terminology reference
└── README.md # Implementation guide
```
## Implementation Workflow
### Step 1: Read the Constitution
```bash
# The constitution contains 11 principles that constrain all implementations
cat constitution.md
```
Key principles to understand:
- **No unsupervised execution**: Every finding requires explicit confirmation
- **Evidence-gated findings**: Claims without evidence don't become findings
- **Reproducibility**: Every finding must be reproducible from its evidence
- **Atomic progress**: Claims are indivisible units of work
- **Fail-safe defaults**: When stuck, escalate or yield—never guess
### Step 2: Install spec-kit
```bash
# In your project directory
npm install -g @github/spec-kit
# or follow spec-kit installation for your coding agent
# Initialize in your project
cd your-security-eval-project
speckit init
```
This creates `.specify/` directory for spec-driven development.
### Step 3: Register the Constitution
```bash
# Copy constitution into spec-kit memory
cp path/to/foundry-security-spec/constitution.md .specify/memory/constitution.md
# Register it with your agent
/speckit.constitution
# Select "adopt existing constitution" when prompted
```
### Step 4: Seed Your Specification
```bash
# Create specs directory
mkdir -p specs/001-foundry
# Copy seed specification
cp path/to/foundry-security-spec/spec.md specs/001-foundry/spec.md
```
### Step 5: Clarify for Your Environment
```bash
# Run clarification workflow
/speckit.clarify
```
Answer questions in these categories:
**Identity & Scope:**
```
Q: What is your system name?
A: acme-security-eval
Q: Does "authorized evaluation with source access" hold?
A: yes
Q: Merge, split, or keep the 8 core roles as-is?
A: keep as-is for first implementation
```
**Integration Choices:**
```
Q: Version control system?
A: GitLab self-hosted at https://gitlab.internal
Q: Issue tracker?
A: Jira at https://jira.acme.com
Q: LLM provider?
A: OpenAI via internal gateway at https://llm.internal/v1
Q: Datastore?
A: PostgreSQL
Q: Isolation runtime?
A: Docker containers with network isolation
Q: Deployment target?
A: Kubernetes cluster
```
**Policy Choices:**
```
Q: Severity taxonomy?
A: Critical/High/Medium/Low matching our existing CVE scale
Q: Surface needs-review findings?
A: No, validator rejects inconclusive findings
Q: Label naming convention?
A: foundry:role/name format
```
**Extension Scope (recommend NO for first build):**
```
Q: Include Attack-Mapper role?
A: no
Q: Include Regression-Tracker role?
A: no
Q: Include Compliance-Mapper role?
A: no
Q: Include Impact-Assessor role?
A: no
Q: Include Remediation-Drafter role?
A: no
```
### Step 6: Generate Your Specification
```bash
# Harden clarified spec
/speckit.specify
# Check for remaining clarifications
/speckit.clarify
# Repeat until no markers remain
```
Your `specs/001-foundry/spec.md` now contains YOUR specification with decisions filled in.
### Step 7: Implement
```bash
# Generate technical design
/speckit.plan
# Generate task backlog
/speckit.tasks
# Start implementation
/speckit.implement
```
## Agent Role Implementation Examples
### Orchestrator Pattern
```python
# orchestrator.py
import asyncio
from typing import List, Dict
from datastore import FindingStore, ClaimStore
from agents import Planner, Detector, Explorer, Validator
class Orchestrator:
def __init__(self,
llm_client,
finding_store: FindingStore,
claim_store: ClaimStore,
budget_manager):
self.llm = llm_client
self.findings = finding_store
self.claims = claim_store
self.budget = budget_manager
# Initialize agent roles
self.planner = Planner(llm_client)
self.detector = Detector(llm_client, rules_corpus)
self.explorer = Explorer(llm_client)
self.validator = Validator(llm_client)
async def evaluate(self, target_repo: str) -> Dict:
"""Run complete evaluation coordinating all agents."""
# FR-001: Orchestrator creates evaluation record
eval_id = await self.findings.create_evaluation(
target=target_repo,
status="running"
)
try:
# FR-010: Planner creates work plan
plan = await self.planner.create_plan(target_repo)
await self.claims.store_plan(eval_id, plan)
# FR-020: Distribute work to detection and exploration
detection_task = asyncio.create_task(
self.run_detection(eval_id, plan)
)
exploration_task = asyncio.create_task(
self.run_exploration(eval_id, plan)
)
# FR-005: Monitor heartbeats and budgets
monitor_task = asyncio.create_task(
self.monitor_health(eval_id)
)
# Wait for completion
await asyncio.gather(
detection_task,
exploration_task,
monitor_task
)
# FR-006: Check coverage gate before completion
coverage = await self.calculate_coverage(eval_id)
if coverage < plan.required_coverage:
raise InsufficientCoverageError(
f"Coverage {coverage}% < required {plan.required_coverage}%"
)
# Mark evaluation complete
await self.findings.update_evaluation(
eval_id,
status="complete",
coverage=coverage
)
return {
"eval_id": eval_id,
"status": "complete",
"findings": await self.findings.count(eval_id),
"coverage": coverage
}
except Exception as e:
# FR-008: Fail-safe: mark evaluation failed
await self.findings.update_evaluation(
eval_id,
status="failed",
error=str(e)
)
raise
async def monitor_health(self, eval_id: str):
"""Monitor agent heartbeats and budgets."""
while True:
await asyncio.sleep(30)
# FR-007: Check heartbeats
stalled = await self.claims.find_stalled_claims(
eval_id,
heartbeat_threshold=300 # 5 minutes
)
for claim in stalled:
# FR-007: Auto-block stalled claims
await self.claims.block_claim(
claim.id,
reason="heartbeat_timeout"
)
# FR-009: Check budget exhaustion
if await self.budget.is_exhausted(eval_id):
await self.findings.update_evaluation(
eval_id,
status="budget_exhausted"
)
break
```
### Detector with CodeGuard Rules
```python
# detector.py
from typing import List
from codeguard import RuleEngine, Finding as CodeGuardFinding
from models import Claim, Finding
class Detector:
def __init__(self, llm_client, rules_corpus_path: str):
self.llm = llm_client
# FR-030: Load CodeGuard rules
self.rule_engine = RuleEngine.load(rules_corpus_path)
async def process_claim(self, claim: Claim) -> List[Finding]:
"""Apply detection rules to a code claim."""
# FR-031: Extract relevant code from claim
code_units = await self.extract_code_units(claim)
findings = []
for unit in code_units:
# FR-032: Run rule engine
rule_hits = await self.rule_engine.evaluate(
code=unit.content,
context=unit.context,
language=unit.language
)
for hit in rule_hits:
# FR-033: Convert rule hit to finding
finding = Finding(
claim_id=claim.id,
rule_id=hit.rule_id,
severity=hit.severity,
weakness_id=hit.cwe_id,
location=hit.location,
evidence={
"rule_match": hit.matched_pattern,
"code_snippet": unit.content,
"line_range": hit.line_range
},
verdict="confirmed", # Rules are deterministic
status="validated"
)
findings.append(finding)
# FR-034: Record coverage
await self.record_coverage(claim.id, unit.path)
return findings
async def extract_code_units(self, claim: Claim):
"""Use LLM to identify relevant code units in claim scope."""
prompt = f"""
Claim: {claim.description}
Scope: {claim.scope}
Identify all code units (functions, methods, classes) that should be
evaluated for security issues related to this claim.
Return as JSON array with: path, name, start_line, end_line
"""
response = await self.llm.complete(prompt)
return parse_code_units(response)
```
### Explorer for Novel Issues
```python
# explorer.py
import asyncio
from typing import List, Optional
from models import Claim, Finding, RuleGap
class Explorer:
def __init__(self, llm_client, sandbox_runtime):
self.llm = llm_client
self.sandbox = sandbox_runtime
async def investigate_claim(self, claim: Claim) -> List[Finding]:
"""Creative exploration beyond static rules."""
findings = []
# FR-040: Generate investigation hypotheses
hypotheses = await self.generate_hypotheses(claim)
for hypothesis in hypotheses:
# FR-041: Execute in isolated sandbox
async with self.sandbox.session() as session:
result = await self.test_hypothesis(
session,
hypothesis,
claim
)
if result.is_vulnerability:
# FR-042: Evidence-gated finding
if not result.has_reproduction:
# Don't create finding without evidence
continue
finding = Finding(
claim_id=claim.id,
severity=result.severity,
weakness_id=result.weakness_id,
description=result.description,
evidence=result.evidence,
verdict="needs-validation",
status="pending"
)
findings.append(finding)
# FR-043: Check if rules missed this
if await self.should_have_detected(finding):
await self.record_rule_gap(finding)
return findings
async def generate_hypotheses(self, claim: Claim) -> List[Dict]:
"""Use LLM to generate creative test hypotheses."""
prompt = f"""
You are exploring code for security issues that static rules may miss.
Claim: {claim.description}
Code scope: {claim.scope}
Generate 3-5 security hypotheses to test:
- Focus on logic bugs, state confusion, race conditions
- Consider what rules can't express (context-dependent issues)
- Prioritize high-impact scenarios
For each hypothesis provide:
- What to test
- Why it might be vulnerable
- How to reproduce if vulnerable
Return as JSON array.
"""
response = await self.llm.complete(
prompt,
temperature=0.7 # Higher for creative exploration
)
return parse_hypotheses(response)
async def record_rule_gap(self, finding: Finding):
"""Record that rules failed to detect this issue."""
gap = RuleGap(
finding_id=finding.id,
weakness_id=finding.weakness_id,
pattern=finding.evidence.get("vulnerable_pattern"),
reason="explorer_found_missed_by_detector",
suggested_rule=await self.draft_rule(finding)
)
# FR-044: Feed into rule corpus improvement
await self.rule_gaps.store(gap)
```
### Validator for Finding Confirmation
```python
# validator.py
from models import Finding, ValidationResult
class Validator:
def __init__(self, llm_client, sandbox_runtime):
self.llm = llm_client
self.sandbox = sandbox_runtime
async def validate_finding(self, finding: Finding) -> ValidationResult:
"""Reproduce and confirm finding from evidence."""
# FR-050: Check evidence completeness
if not self.has_sufficient_evidence(finding):
return ValidationResult(
verdict="rejected",
reason="insufficient_evidence"
)
# FR-051: Attempt reproduction
async with self.sandbox.session() as session:
reproduced = await self.reproduce_issue(
session,
finding.evidence
)
if not reproduced:
return ValidationResult(
verdict="rejected",
reason="not_reproducible"
)
# FR-052: Verify severity assessment
actual_severity = await self.assess_severity(
session,
finding
)
if actual_severity != finding.severity:
finding.severity = actual_severity
finding.evidence["severity_adjustment"] = {
"original": finding.severity,
"validated": actual_severity
}
# FR-053: Generate fingerprint for deduplication
fingerprint = await self.generate_fingerprint(finding)
return ValidationResult(
verdict="confirmed",
fingerprint=fingerprint,
severity=actual_severity,
reproduction_evidence=session.get_transcript()
)
def has_sufficient_evidence(self, finding: Finding) -> bool:
"""Check if finding has required evidence."""
required = ["location", "description"]
if finding.severity in ["critical", "high"]:
required.extend(["reproduction_steps", "impact"])
return all(k in finding.evidence for k in required)
async def generate_fingerprint(self, finding: Finding) -> str:
"""Create stable fingerprint for deduplication."""
# FR-054: Fingerprint combines weakness + location + root cause
components = [
finding.weakness_id,
finding.location.get("file_path"),
finding.location.get("function_name"),
finding.evidence.get("root_cause_pattern")
]
Auf GitHub ansehen