Skip to main content

foundry-security-spec

Implement Cisco's Foundry specification for agentic AI security evaluation systems with multi-agent architecture

Quellinformationen

Repository
reason-machines/security-skills
Letzte Quellaktivität
22. Mai 2026 um 22:52
Erkannte Sprache von SKILL.md
Englisch
Sterne
12
Forks
1

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
foundry-security-spec
description
Implement Cisco's Foundry specification for agentic AI security evaluation systems with multi-agent architecture
triggers
["implement foundry security evaluation","set up agentic security testing","create foundry spec system","build security evaluation agents","configure foundry security roles","design ai vulnerability discovery","implement foundry detector agents","create security finding workflow"]
# Foundry Security Spec > Skill by [ara.so](https://ara.so) — Security Skills collection. Foundry is an open specification from Cisco for building agentic AI security evaluation systems. It defines a multi-agent architecture with 8 core roles and 5 extension roles that coordinate to discover, validate, and report security findings. This is NOT a tool to install—it's a blueprint for building your own security evaluation system. ## Core Concepts Foundry provides: - **Architecture**: 8 core agent roles (Orchestrator, Planner, Navigator, Detector, Explorer, Validator, Investigator, Publisher) - **Finding Lifecycle**: States, verdicts, evidence gates, fingerprinting - **Coordination Model**: Atomic claims, heartbeat liveness, auto-blocking - **Governance**: Sandboxing, budgets, yield-gated auto-stop, coverage gates - **Detection-to-Prevention Flywheel**: Rules catch known issues, explorers find new ones, gaps become new rules Works with CodeGuard rule format for portable detection rules that transfer between evaluation and prevention. ## Repository Structure ``` foundry-security-spec/ ├── spec.md # Main specification (~130 functional requirements) ├── constitution.md # 11 inviolable principles ├── GLOSSARY.md # Terminology reference └── README.md # Implementation guide ``` ## Implementation Workflow ### Step 1: Read the Constitution ```bash # The constitution contains 11 principles that constrain all implementations cat constitution.md ``` Key principles to understand: - **No unsupervised execution**: Every finding requires explicit confirmation - **Evidence-gated findings**: Claims without evidence don't become findings - **Reproducibility**: Every finding must be reproducible from its evidence - **Atomic progress**: Claims are indivisible units of work - **Fail-safe defaults**: When stuck, escalate or yield—never guess ### Step 2: Install spec-kit ```bash # In your project directory npm install -g @github/spec-kit # or follow spec-kit installation for your coding agent # Initialize in your project cd your-security-eval-project speckit init ``` This creates `.specify/` directory for spec-driven development. ### Step 3: Register the Constitution ```bash # Copy constitution into spec-kit memory cp path/to/foundry-security-spec/constitution.md .specify/memory/constitution.md # Register it with your agent /speckit.constitution # Select "adopt existing constitution" when prompted ``` ### Step 4: Seed Your Specification ```bash # Create specs directory mkdir -p specs/001-foundry # Copy seed specification cp path/to/foundry-security-spec/spec.md specs/001-foundry/spec.md ``` ### Step 5: Clarify for Your Environment ```bash # Run clarification workflow /speckit.clarify ``` Answer questions in these categories: **Identity & Scope:** ``` Q: What is your system name? A: acme-security-eval Q: Does "authorized evaluation with source access" hold? A: yes Q: Merge, split, or keep the 8 core roles as-is? A: keep as-is for first implementation ``` **Integration Choices:** ``` Q: Version control system? A: GitLab self-hosted at https://gitlab.internal Q: Issue tracker? A: Jira at https://jira.acme.com Q: LLM provider? A: OpenAI via internal gateway at https://llm.internal/v1 Q: Datastore? A: PostgreSQL Q: Isolation runtime? A: Docker containers with network isolation Q: Deployment target? A: Kubernetes cluster ``` **Policy Choices:** ``` Q: Severity taxonomy? A: Critical/High/Medium/Low matching our existing CVE scale Q: Surface needs-review findings? A: No, validator rejects inconclusive findings Q: Label naming convention? A: foundry:role/name format ``` **Extension Scope (recommend NO for first build):** ``` Q: Include Attack-Mapper role? A: no Q: Include Regression-Tracker role? A: no Q: Include Compliance-Mapper role? A: no Q: Include Impact-Assessor role? A: no Q: Include Remediation-Drafter role? A: no ``` ### Step 6: Generate Your Specification ```bash # Harden clarified spec /speckit.specify # Check for remaining clarifications /speckit.clarify # Repeat until no markers remain ``` Your `specs/001-foundry/spec.md` now contains YOUR specification with decisions filled in. ### Step 7: Implement ```bash # Generate technical design /speckit.plan # Generate task backlog /speckit.tasks # Start implementation /speckit.implement ``` ## Agent Role Implementation Examples ### Orchestrator Pattern ```python # orchestrator.py import asyncio from typing import List, Dict from datastore import FindingStore, ClaimStore from agents import Planner, Detector, Explorer, Validator class Orchestrator: def __init__(self, llm_client, finding_store: FindingStore, claim_store: ClaimStore, budget_manager): self.llm = llm_client self.findings = finding_store self.claims = claim_store self.budget = budget_manager # Initialize agent roles self.planner = Planner(llm_client) self.detector = Detector(llm_client, rules_corpus) self.explorer = Explorer(llm_client) self.validator = Validator(llm_client) async def evaluate(self, target_repo: str) -> Dict: """Run complete evaluation coordinating all agents.""" # FR-001: Orchestrator creates evaluation record eval_id = await self.findings.create_evaluation( target=target_repo, status="running" ) try: # FR-010: Planner creates work plan plan = await self.planner.create_plan(target_repo) await self.claims.store_plan(eval_id, plan) # FR-020: Distribute work to detection and exploration detection_task = asyncio.create_task( self.run_detection(eval_id, plan) ) exploration_task = asyncio.create_task( self.run_exploration(eval_id, plan) ) # FR-005: Monitor heartbeats and budgets monitor_task = asyncio.create_task( self.monitor_health(eval_id) ) # Wait for completion await asyncio.gather( detection_task, exploration_task, monitor_task ) # FR-006: Check coverage gate before completion coverage = await self.calculate_coverage(eval_id) if coverage < plan.required_coverage: raise InsufficientCoverageError( f"Coverage {coverage}% < required {plan.required_coverage}%" ) # Mark evaluation complete await self.findings.update_evaluation( eval_id, status="complete", coverage=coverage ) return { "eval_id": eval_id, "status": "complete", "findings": await self.findings.count(eval_id), "coverage": coverage } except Exception as e: # FR-008: Fail-safe: mark evaluation failed await self.findings.update_evaluation( eval_id, status="failed", error=str(e) ) raise async def monitor_health(self, eval_id: str): """Monitor agent heartbeats and budgets.""" while True: await asyncio.sleep(30) # FR-007: Check heartbeats stalled = await self.claims.find_stalled_claims( eval_id, heartbeat_threshold=300 # 5 minutes ) for claim in stalled: # FR-007: Auto-block stalled claims await self.claims.block_claim( claim.id, reason="heartbeat_timeout" ) # FR-009: Check budget exhaustion if await self.budget.is_exhausted(eval_id): await self.findings.update_evaluation( eval_id, status="budget_exhausted" ) break ``` ### Detector with CodeGuard Rules ```python # detector.py from typing import List from codeguard import RuleEngine, Finding as CodeGuardFinding from models import Claim, Finding class Detector: def __init__(self, llm_client, rules_corpus_path: str): self.llm = llm_client # FR-030: Load CodeGuard rules self.rule_engine = RuleEngine.load(rules_corpus_path) async def process_claim(self, claim: Claim) -> List[Finding]: """Apply detection rules to a code claim.""" # FR-031: Extract relevant code from claim code_units = await self.extract_code_units(claim) findings = [] for unit in code_units: # FR-032: Run rule engine rule_hits = await self.rule_engine.evaluate( code=unit.content, context=unit.context, language=unit.language ) for hit in rule_hits: # FR-033: Convert rule hit to finding finding = Finding( claim_id=claim.id, rule_id=hit.rule_id, severity=hit.severity, weakness_id=hit.cwe_id, location=hit.location, evidence={ "rule_match": hit.matched_pattern, "code_snippet": unit.content, "line_range": hit.line_range }, verdict="confirmed", # Rules are deterministic status="validated" ) findings.append(finding) # FR-034: Record coverage await self.record_coverage(claim.id, unit.path) return findings async def extract_code_units(self, claim: Claim): """Use LLM to identify relevant code units in claim scope.""" prompt = f""" Claim: {claim.description} Scope: {claim.scope} Identify all code units (functions, methods, classes) that should be evaluated for security issues related to this claim. Return as JSON array with: path, name, start_line, end_line """ response = await self.llm.complete(prompt) return parse_code_units(response) ``` ### Explorer for Novel Issues ```python # explorer.py import asyncio from typing import List, Optional from models import Claim, Finding, RuleGap class Explorer: def __init__(self, llm_client, sandbox_runtime): self.llm = llm_client self.sandbox = sandbox_runtime async def investigate_claim(self, claim: Claim) -> List[Finding]: """Creative exploration beyond static rules.""" findings = [] # FR-040: Generate investigation hypotheses hypotheses = await self.generate_hypotheses(claim) for hypothesis in hypotheses: # FR-041: Execute in isolated sandbox async with self.sandbox.session() as session: result = await self.test_hypothesis( session, hypothesis, claim ) if result.is_vulnerability: # FR-042: Evidence-gated finding if not result.has_reproduction: # Don't create finding without evidence continue finding = Finding( claim_id=claim.id, severity=result.severity, weakness_id=result.weakness_id, description=result.description, evidence=result.evidence, verdict="needs-validation", status="pending" ) findings.append(finding) # FR-043: Check if rules missed this if await self.should_have_detected(finding): await self.record_rule_gap(finding) return findings async def generate_hypotheses(self, claim: Claim) -> List[Dict]: """Use LLM to generate creative test hypotheses.""" prompt = f""" You are exploring code for security issues that static rules may miss. Claim: {claim.description} Code scope: {claim.scope} Generate 3-5 security hypotheses to test: - Focus on logic bugs, state confusion, race conditions - Consider what rules can't express (context-dependent issues) - Prioritize high-impact scenarios For each hypothesis provide: - What to test - Why it might be vulnerable - How to reproduce if vulnerable Return as JSON array. """ response = await self.llm.complete( prompt, temperature=0.7 # Higher for creative exploration ) return parse_hypotheses(response) async def record_rule_gap(self, finding: Finding): """Record that rules failed to detect this issue.""" gap = RuleGap( finding_id=finding.id, weakness_id=finding.weakness_id, pattern=finding.evidence.get("vulnerable_pattern"), reason="explorer_found_missed_by_detector", suggested_rule=await self.draft_rule(finding) ) # FR-044: Feed into rule corpus improvement await self.rule_gaps.store(gap) ``` ### Validator for Finding Confirmation ```python # validator.py from models import Finding, ValidationResult class Validator: def __init__(self, llm_client, sandbox_runtime): self.llm = llm_client self.sandbox = sandbox_runtime async def validate_finding(self, finding: Finding) -> ValidationResult: """Reproduce and confirm finding from evidence.""" # FR-050: Check evidence completeness if not self.has_sufficient_evidence(finding): return ValidationResult( verdict="rejected", reason="insufficient_evidence" ) # FR-051: Attempt reproduction async with self.sandbox.session() as session: reproduced = await self.reproduce_issue( session, finding.evidence ) if not reproduced: return ValidationResult( verdict="rejected", reason="not_reproducible" ) # FR-052: Verify severity assessment actual_severity = await self.assess_severity( session, finding ) if actual_severity != finding.severity: finding.severity = actual_severity finding.evidence["severity_adjustment"] = { "original": finding.severity, "validated": actual_severity } # FR-053: Generate fingerprint for deduplication fingerprint = await self.generate_fingerprint(finding) return ValidationResult( verdict="confirmed", fingerprint=fingerprint, severity=actual_severity, reproduction_evidence=session.get_transcript() ) def has_sufficient_evidence(self, finding: Finding) -> bool: """Check if finding has required evidence.""" required = ["location", "description"] if finding.severity in ["critical", "high"]: required.extend(["reproduction_steps", "impact"]) return all(k in finding.evidence for k in required) async def generate_fingerprint(self, finding: Finding) -> str: """Create stable fingerprint for deduplication.""" # FR-054: Fingerprint combines weakness + location + root cause components = [ finding.weakness_id, finding.location.get("file_path"), finding.location.get("function_name"), finding.evidence.get("root_cause_pattern") ]
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen