Skip to main content

create-evidence-case

Build structured evidence cases using Claims-Arguments-Evidence (CAE) trees. AGENT-DRIVEN composable orchestrator: the agent decomposes the question, calls existing skills for data collection and verification, then DECIDES the verdict. Python runner.py is a thin data collector and persistence layer.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
grahama1970/agent-skills
آخر نشاط في المصدر
٨ أغسطس ٢٠٢٦ في ١٦:٣٢
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٥
التفرعات
٢

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
86 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
create-evidence-case
description
Build structured evidence cases using Claims-Arguments-Evidence (CAE) trees. AGENT-DRIVEN composable orchestrator: the agent decomposes the question, calls existing skills for data collection and verification, then DECIDES the verdict. Python runner.py is a thin data collector and persistence layer.
allowed-tools
["Bash","Read","Grep","Glob","Write","Edit"]
triggers
["create-evidence-case","evidence case","build evidence","verify claim","evidence tree"]
metadata
{"short-description":"Composable CAE evidence orchestrator","version":"4.3.0"}
provides
["evidence-case-creation","uct-strategy-selection","evidence-tree-persistence","claim-verdict-grading","assurance-case-output","oscal-export","qra-candidate-quarantine","human-review-interview"]
composes
["memory","extract-entities","assistant","taxonomy","lean4-prove","dogpile","edge-verifier","cmmc-assessor","create-gsn-diagram","export-oscal","create-figure","task-monitor","interview","prompt-lab","analytics","agentic-evals"]
taxonomy
["verification","evidence","quality-assurance"]
disciplines
["compliance-security","evaluation-quality"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT. **User-facing documentation:** See [COMPLETE_EXAMPLE.md](COMPLETE_EXAMPLE.md) for compliance officer workflow, CAE tree examples, and the tool's role as a research assistant (not compliance oracle). # /create-evidence-case v4.3 — Composable Orchestrator Build structured **Claims-Arguments-Evidence (CAE)** trees. This is an **agent-driven** skill — you (the agent) orchestrate existing skills, reason about results, and make all judgment calls. The Python code is a thin data collector. ## PROVIDER BOUNDARY Do not call `$scillm`, `/scillm`, `http://localhost:4001`, `/v1/chat/completions`, or `/v1/scillm/*` directly from a `/create-evidence-case` project-agent workflow. Provider/model execution is owned by Tau. If model rendering, batch review, shadow collection, or provider delegation is required, express it as a Tau DAG contract, Tau skill node, or the owning skill's Tau-mediated runtime and consume the resulting receipts. Direct SciLLM use is allowed only for Tau/SciLLM maintenance or when the human explicitly asks to operate SciLLM itself. A CAE verdict must remain grounded in Memory recall, `/extract-entities` proof packets, same-technique checks, and deterministic gates; raw model output is never evidence. ## INPUT/OUTPUT MODEL (v4.2) **Input**: Any control from `sparta_controls` (SPARTA, NIST, CWE, CAPEC) OR a raw question. **Output**: Always a QRA document in `sparta_qra`, regardless of input source. | Input Source | Example | Output | |--------------|---------|--------| | SPARTA control | SA-1 | QRA with `source_framework: "SPARTA"` | | NIST control | IA-5 | QRA with `source_framework: "NIST-800-53"` | | CWE | CWE-287 | QRA with `source_framework: "CWE"` | | Raw question | "What is CWE-287?" | QRA with extracted `source_control_id` | ### QRA Output Tiers Not every QRA has evidence. The `evidence_case` field being `null` is a **deterministic signal** (no crosswalk edges exist), not a failure. | Tier | evidence_case | Use Case | |------|---------------|----------| | **Informational** | null | Explanation, lookup (no crosswalk edges) | | **Grounded** | `{chains: [...], formal_proof: null}` | Compliance, audit | | **Verified** | `{chains: [...], formal_proof: {success: true}}` | Safety-critical, formal methods | All fields (chains, formal_proof, sacm_ref) live **inside** evidence_case. One null check. ### Entity Extraction → Evidence Case Flow ``` "What is CWE-287?" │ ▼ /extract-entities → "CWE-287" ✓ │ ▼ /create-evidence-case ├── Query crosswalk edges in sparta_relationships ├── Has edges → evidence_case: {chains: [...]} └── No edges → evidence_case: null (valid output) │ ▼ QRA created in sparta_qra ├── question: "What is CWE-287?" ├── requirement: CWE-287 description ├── answer: Explanation text ├── source_framework: "CWE" ├── source_control_id: "CWE-287" └── evidence_case: {...} or null # SELF-CONTAINED: includes chains, proof, sacm_ref ``` ### Self-Contained QRA Schema (v4.3) Every QRA embeds all metadata for downstream consumption. **evidence_case is fully self-contained** — formal proofs and SACM references live INSIDE it, not at top level. This ensures traceability in ArangoDB and portability for OSCAL/SACM export. ```python { # Core QRA fields "question": str, "requirement": str, "answer": str, "source_framework": str, # SPARTA, NIST-800-53, CWE, CAPEC "source_control_id": str, # SA-1, IA-5, CWE-287 "relationship_id": str, # Links persona variations "expertise": str, # expert, layperson, pm, compliance "difficulty": str, # single_hop, multi_hop, synthesis # Taxonomy (Mind tags for SPARTA scope) "mind": list[str], # ["Detect", "Harden"] # Evidence case — SELF-CONTAINED (null if no crosswalk edges) # All verification evidence lives here: chains, proof, SACM ref "evidence_case": { "chains": [...], # Crosswalk paths (source → hops → target) "confidence": float, # 0.0-1.0 combined confidence "methods": list[str], # ["direct", "mitre_chain", "nist_nvd"] # Formal proof (null if not verified) — INSIDE evidence_case "formal_proof": { "success": bool, "code": str, # Lean4 source code "attempts": int, "errors": list | None, "proved_at": int | None, } | None, # SACM reference (null if not exported) — INSIDE evidence_case "sacm_ref": { "gid": str, # SACM Goal node GID (lookup key) "xml_snippet": str, # First 500 chars (preview) "generated_at": int, } | None, } | None, } ``` **Why self-contained?** ISO 26262 and DO-178C refinement checking patterns bundle verification evidence with the claim. Separate collections would require joins, lose portability, and break single-document export to OSCAL/SACM XML. ### Cached Evidence Case in QRAs (v4.3.1) Every QRA has a pre-computed `evidence_case` field containing extracted entities with **character spans** for UI inline highlighting. This is populated by the `evidence_case_backfill` pipeline (not query-time extraction). ```python { "evidence_case": { # Core extracted data (from /extract-entities) "question_text": str, # First 500 chars of question "control_ids": list[str], # ["IA-0006", "CWE-287"] "glossary": [ # Snapshot copied from /extract-entities.glossary { "id": "IA-0006", "name": "Authentication Mechanisms", "framework": "SPARTA", "type": "countermeasure", "description": "..." } ], "crosswalk_chains": [...], # Framework-to-framework links "resolved_entities": [...], # Terms that resolved to controls # --- SPANS FOR UI HIGHLIGHTING (SPARTA Explorer Chat) --- "spans": [ { "text": "IA-0006", # The matched text "span": [9, 16], # [start, end] character positions "kind": "control_id", # Type: control_id, aerospace_term, phrase "framework": "SPARTA", # Only for control_id kind "name": "Authentication Mechanisms", "grounded_to_framework": True }, { "text": "CWE-287", "span": [33, 40], "kind": "control_id", "framework": "CWE", "name": "Improper Authentication", "grounded_to_framework": True } ], "review_status": "auto", # auto, approved, rejected "extracted_at": str # ISO timestamp } } ``` ### SPARTA Explorer Chat: Inline Evidence Highlighting The Chat agent uses `evidence_case.spans` to render inline highlights on question text: ```tsx // React component example function QuestionWithHighlights({ question, spans }) { // Sort spans by start position const sorted = [...spans].sort((a, b) => a.span[0] - b.span[0]); let result = []; let lastEnd = 0; for (const s of sorted) { const [start, end] = s.span; // Text before this span if (start > lastEnd) { result.push(<span key={`t-${lastEnd}`}>{question.slice(lastEnd, start)}</span>); } // The highlighted entity const color = s.kind === 'control_id' ? FRAMEWORK_COLORS[s.framework] : 'yellow'; result.push( <Tooltip key={`e-${start}`} content={`${s.name} (${s.framework})`}> <span className={`bg-${color}-200`}>{s.text}</span> </Tooltip> ); lastEnd = end; } // Remaining text if (lastEnd < question.length) { result.push(<span key="tail">{question.slice(lastEnd)}</span>); } return <p>{result}</p>; } ``` **Span kinds and colors:** | Kind | Color | Example | |------|-------|---------| | `control_id` | Framework-specific (SPARTA=blue, CWE=orange, NIST=green) | IA-0006, CWE-287 | | `aerospace_term` | Purple | FPGA, TT&C, crosslink | | `phrase` | Yellow | supply chain attack | **Framework colors (NVIS standard):** | Framework | Color | Hex | |-----------|-------|-----| | SPARTA | Blue | `#3B82F6` | | CWE | Orange | `#F97316` | | NIST | Green | `#22C55E` | | CAPEC | Red | `#EF4444` | | ATT&CK | Purple | `#A855F7` | ### Cached vs Query-Time Evidence | Field | When Populated | Use | |-------|----------------|-----| | `evidence_case.spans` | Backfill (pre-computed) | UI highlighting | | `evidence_case.glossary` | Backfill | Cached `/extract-entities.glossary` metadata | | `cae_tree` (in endpoint response) | Query-time | Full CAE tree with traceability | | `entity_context` (in endpoint response) | Query-time | Compact `/extract-entities` agent view with authoritative `proof_packet.assertions`, `nodes`, and candidate-only boundaries | The cached `evidence_case` enables fast rendering without query-time entity extraction. Glossary semantics are owned by `/extract-entities`; `/create-evidence-case` copies that glossary into the evidence case and must not independently infer control names, definitions, or unsupported parenthetical terms. Runtime extraction proof is preserved in `entity_context`; agents should use `entity_context.proof_packet.assertions` for grounding decisions and keep legacy glossary fields for compatibility, prompts, and UI display. The `cae_tree` (requested via `include_cae_tree=True`) provides full CAE analysis with traceability strength and validation gates. ## TRACEABILITY vs VERIFICATION Matching requirements to controls (via `/match/requirement`) provides **traceability** — linking related concepts. But traceability alone is NOT SUFFICIENT for certification. | Level | What it proves | Provides | Sufficient for | |-------|----------------|----------|----------------| | **Matching** | "Related concepts" | Traceability | Gap detection, triage | | **Confidence scoring** | "Probably covers" | Prioritization | Audit prep | | **Formal proof** | "Functionally entails" | Verification | ISO 26262, DO-178C certification | **Why matching is insufficient:** ``` Requirement: "shall implement multi-factor authentication for all users" Control: SV-AC-2 "implement access control mechanisms" Matching confidence: 0.82 - Crosswalk chain exists (CWE → NIST → SPARTA) - QRA evidence mentions both - Semantic similarity is high Problem: Does SV-AC-2 ACTUALLY cover MFA? - Words overlap, but FUNCTIONS might not - Control might be narrower/broader than requirement - Edge cases might not be addressed ``` **Why formal proofs are needed:** ``` Formal proof (Lean4): theorem control_satisfies_requirement : ∀ u : User, implements(SV-AC-2, u) → authenticated(u) ∧ factors(u) ≥ 2 Proof result: success=true Meaning: The control ENTAILS the requirement — not just similar words, but the control is a valid implementation of the requirement. ``` The `evidence_case.formal_proof` field stores this verification evidence. Matching builds the edges; formal proofs prove they're functionally valid. ## EXECUTION MODES | Mode | Engine | When | Latency | |------|--------|------|---------| | **Live** | Project agent (you) | Interactive use | ~10-12s (dominated by `/memory recall`) | | **Batch** | `EvidenceCaseRunner` | Nightly eval, `run_question_bank.py` | ~17s (subprocess classifiers) | **Live mode**: YOU are the engine. Call `/memory recall` for data, do decomposition + entity analysis + same-technique check + verdict in your reasoning. No `/assistant classify` calls needed — you ARE the classifier. `/lean4-prove` Docker compilation is near-instant.
عرض على GitHub
ملف SKILL.md هذا كبير جدا، لذلك يعرض SkillsMP القسم الاول فقط هنا. عرض على GitHub