Skip to main content

create-evidence-case

Build structured evidence cases using Claims-Arguments-Evidence (CAE) trees. AGENT-DRIVEN composable orchestrator: the agent decomposes the question, calls existing skills for data collection and verification, then DECIDES the verdict. Python runner.py is a thin data collector and persistence layer.

Ir a la instalación

Datos de origen

Repositorio
grahama1970/agent-skills
Última actividad en el origen
8 de agosto de 2026 a las 16:32
Idioma detectado de SKILL.md
inglés
Estrellas
5
Forks
2

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Explorador de archivos
86 archivos

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
create-evidence-case
description
Build structured evidence cases using Claims-Arguments-Evidence (CAE) trees. AGENT-DRIVEN composable orchestrator: the agent decomposes the question, calls existing skills for data collection and verification, then DECIDES the verdict. Python runner.py is a thin data collector and persistence layer.
allowed-tools
["Bash","Read","Grep","Glob","Write","Edit"]
triggers
["create-evidence-case","evidence case","build evidence","verify claim","evidence tree"]
metadata
{"short-description":"Composable CAE evidence orchestrator","version":"4.3.0"}
provides
["evidence-case-creation","uct-strategy-selection","evidence-tree-persistence","claim-verdict-grading","assurance-case-output","oscal-export","qra-candidate-quarantine","human-review-interview"]
composes
["memory","extract-entities","assistant","taxonomy","lean4-prove","dogpile","edge-verifier","cmmc-assessor","create-gsn-diagram","export-oscal","create-figure","task-monitor","interview","prompt-lab","analytics","agentic-evals"]
taxonomy
["verification","evidence","quality-assurance"]
disciplines
["compliance-security","evaluation-quality"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT. **User-facing documentation:** See [COMPLETE_EXAMPLE.md](COMPLETE_EXAMPLE.md) for compliance officer workflow, CAE tree examples, and the tool's role as a research assistant (not compliance oracle). # /create-evidence-case v4.3 — Composable Orchestrator Build structured **Claims-Arguments-Evidence (CAE)** trees. This is an **agent-driven** skill — you (the agent) orchestrate existing skills, reason about results, and make all judgment calls. The Python code is a thin data collector. ## PROVIDER BOUNDARY Do not call `$scillm`, `/scillm`, `http://localhost:4001`, `/v1/chat/completions`, or `/v1/scillm/*` directly from a `/create-evidence-case` project-agent workflow. Provider/model execution is owned by Tau. If model rendering, batch review, shadow collection, or provider delegation is required, express it as a Tau DAG contract, Tau skill node, or the owning skill's Tau-mediated runtime and consume the resulting receipts. Direct SciLLM use is allowed only for Tau/SciLLM maintenance or when the human explicitly asks to operate SciLLM itself. A CAE verdict must remain grounded in Memory recall, `/extract-entities` proof packets, same-technique checks, and deterministic gates; raw model output is never evidence. ## INPUT/OUTPUT MODEL (v4.2) **Input**: Any control from `sparta_controls` (SPARTA, NIST, CWE, CAPEC) OR a raw question. **Output**: Always a QRA document in `sparta_qra`, regardless of input source. | Input Source | Example | Output | |--------------|---------|--------| | SPARTA control | SA-1 | QRA with `source_framework: "SPARTA"` | | NIST control | IA-5 | QRA with `source_framework: "NIST-800-53"` | | CWE | CWE-287 | QRA with `source_framework: "CWE"` | | Raw question | "What is CWE-287?" | QRA with extracted `source_control_id` | ### QRA Output Tiers Not every QRA has evidence. The `evidence_case` field being `null` is a **deterministic signal** (no crosswalk edges exist), not a failure. | Tier | evidence_case | Use Case | |------|---------------|----------| | **Informational** | null | Explanation, lookup (no crosswalk edges) | | **Grounded** | `{chains: [...], formal_proof: null}` | Compliance, audit | | **Verified** | `{chains: [...], formal_proof: {success: true}}` | Safety-critical, formal methods | All fields (chains, formal_proof, sacm_ref) live **inside** evidence_case. One null check. ### Entity Extraction → Evidence Case Flow ``` "What is CWE-287?" │ ▼ /extract-entities → "CWE-287" ✓ │ ▼ /create-evidence-case ├── Query crosswalk edges in sparta_relationships ├── Has edges → evidence_case: {chains: [...]} └── No edges → evidence_case: null (valid output) │ ▼ QRA created in sparta_qra ├── question: "What is CWE-287?" ├── requirement: CWE-287 description ├── answer: Explanation text ├── source_framework: "CWE" ├── source_control_id: "CWE-287" └── evidence_case: {...} or null # SELF-CONTAINED: includes chains, proof, sacm_ref ``` ### Self-Contained QRA Schema (v4.3) Every QRA embeds all metadata for downstream consumption. **evidence_case is fully self-contained** — formal proofs and SACM references live INSIDE it, not at top level. This ensures traceability in ArangoDB and portability for OSCAL/SACM export. ```python { # Core QRA fields "question": str, "requirement": str, "answer": str, "source_framework": str, # SPARTA, NIST-800-53, CWE, CAPEC "source_control_id": str, # SA-1, IA-5, CWE-287 "relationship_id": str, # Links persona variations "expertise": str, # expert, layperson, pm, compliance "difficulty": str, # single_hop, multi_hop, synthesis # Taxonomy (Mind tags for SPARTA scope) "mind": list[str], # ["Detect", "Harden"] # Evidence case — SELF-CONTAINED (null if no crosswalk edges) # All verification evidence lives here: chains, proof, SACM ref "evidence_case": { "chains": [...], # Crosswalk paths (source → hops → target) "confidence": float, # 0.0-1.0 combined confidence "methods": list[str], # ["direct", "mitre_chain", "nist_nvd"] # Formal proof (null if not verified) — INSIDE evidence_case "formal_proof": { "success": bool, "code": str, # Lean4 source code "attempts": int, "errors": list | None, "proved_at": int | None, } | None, # SACM reference (null if not exported) — INSIDE evidence_case "sacm_ref": { "gid": str, # SACM Goal node GID (lookup key) "xml_snippet": str, # First 500 chars (preview) "generated_at": int, } | None, } | None, } ``` **Why self-contained?** ISO 26262 and DO-178C refinement checking patterns bundle verification evidence with the claim. Separate collections would require joins, lose portability, and break single-document export to OSCAL/SACM XML. ### Cached Evidence Case in QRAs (v4.3.1) Every QRA has a pre-computed `evidence_case` field containing extracted entities with **character spans** for UI inline highlighting. This is populated by the `evidence_case_backfill` pipeline (not query-time extraction). ```python { "evidence_case": { # Core extracted data (from /extract-entities) "question_text": str, # First 500 chars of question "control_ids": list[str], # ["IA-0006", "CWE-287"] "glossary": [ # Snapshot copied from /extract-entities.glossary { "id": "IA-0006", "name": "Authentication Mechanisms", "framework": "SPARTA", "type": "countermeasure", "description": "..." } ], "crosswalk_chains": [...], # Framework-to-framework links "resolved_entities": [...], # Terms that resolved to controls # --- SPANS FOR UI HIGHLIGHTING (SPARTA Explorer Chat) --- "spans": [ { "text": "IA-0006", # The matched text "span": [9, 16], # [start, end] character positions "kind": "control_id", # Type: control_id, aerospace_term, phrase "framework": "SPARTA", # Only for control_id kind "name": "Authentication Mechanisms", "grounded_to_framework": True }, { "text": "CWE-287", "span": [33, 40], "kind": "control_id", "framework": "CWE", "name": "Improper Authentication", "grounded_to_framework": True } ], "review_status": "auto", # auto, approved, rejected "extracted_at": str # ISO timestamp } } ``` ### SPARTA Explorer Chat: Inline Evidence Highlighting The Chat agent uses `evidence_case.spans` to render inline highlights on question text: ```tsx // React component example function QuestionWithHighlights({ question, spans }) { // Sort spans by start position const sorted = [...spans].sort((a, b) => a.span[0] - b.span[0]); let result = []; let lastEnd = 0; for (const s of sorted) { const [start, end] = s.span; // Text before this span if (start > lastEnd) { result.push(<span key={`t-${lastEnd}`}>{question.slice(lastEnd, start)}</span>); } // The highlighted entity const color = s.kind === 'control_id' ? FRAMEWORK_COLORS[s.framework] : 'yellow'; result.push( <Tooltip key={`e-${start}`} content={`${s.name} (${s.framework})`}> <span className={`bg-${color}-200`}>{s.text}</span> </Tooltip> ); lastEnd = end; } // Remaining text if (lastEnd < question.length) { result.push(<span key="tail">{question.slice(lastEnd)}</span>); } return <p>{result}</p>; } ``` **Span kinds and colors:** | Kind | Color | Example | |------|-------|---------| | `control_id` | Framework-specific (SPARTA=blue, CWE=orange, NIST=green) | IA-0006, CWE-287 | | `aerospace_term` | Purple | FPGA, TT&C, crosslink | | `phrase` | Yellow | supply chain attack | **Framework colors (NVIS standard):** | Framework | Color | Hex | |-----------|-------|-----| | SPARTA | Blue | `#3B82F6` | | CWE | Orange | `#F97316` | | NIST | Green | `#22C55E` | | CAPEC | Red | `#EF4444` | | ATT&CK | Purple | `#A855F7` | ### Cached vs Query-Time Evidence | Field | When Populated | Use | |-------|----------------|-----| | `evidence_case.spans` | Backfill (pre-computed) | UI highlighting | | `evidence_case.glossary` | Backfill | Cached `/extract-entities.glossary` metadata | | `cae_tree` (in endpoint response) | Query-time | Full CAE tree with traceability | | `entity_context` (in endpoint response) | Query-time | Compact `/extract-entities` agent view with authoritative `proof_packet.assertions`, `nodes`, and candidate-only boundaries | The cached `evidence_case` enables fast rendering without query-time entity extraction. Glossary semantics are owned by `/extract-entities`; `/create-evidence-case` copies that glossary into the evidence case and must not independently infer control names, definitions, or unsupported parenthetical terms. Runtime extraction proof is preserved in `entity_context`; agents should use `entity_context.proof_packet.assertions` for grounding decisions and keep legacy glossary fields for compatibility, prompts, and UI display. The `cae_tree` (requested via `include_cae_tree=True`) provides full CAE analysis with traceability strength and validation gates. ## TRACEABILITY vs VERIFICATION Matching requirements to controls (via `/match/requirement`) provides **traceability** — linking related concepts. But traceability alone is NOT SUFFICIENT for certification. | Level | What it proves | Provides | Sufficient for | |-------|----------------|----------|----------------| | **Matching** | "Related concepts" | Traceability | Gap detection, triage | | **Confidence scoring** | "Probably covers" | Prioritization | Audit prep | | **Formal proof** | "Functionally entails" | Verification | ISO 26262, DO-178C certification | **Why matching is insufficient:** ``` Requirement: "shall implement multi-factor authentication for all users" Control: SV-AC-2 "implement access control mechanisms" Matching confidence: 0.82 - Crosswalk chain exists (CWE → NIST → SPARTA) - QRA evidence mentions both - Semantic similarity is high Problem: Does SV-AC-2 ACTUALLY cover MFA? - Words overlap, but FUNCTIONS might not - Control might be narrower/broader than requirement - Edge cases might not be addressed ``` **Why formal proofs are needed:** ``` Formal proof (Lean4): theorem control_satisfies_requirement : ∀ u : User, implements(SV-AC-2, u) → authenticated(u) ∧ factors(u) ≥ 2 Proof result: success=true Meaning: The control ENTAILS the requirement — not just similar words, but the control is a valid implementation of the requirement. ``` The `evidence_case.formal_proof` field stores this verification evidence. Matching builds the edges; formal proofs prove they're functionally valid. ## EXECUTION MODES | Mode | Engine | When | Latency | |------|--------|------|---------| | **Live** | Project agent (you) | Interactive use | ~10-12s (dominated by `/memory recall`) | | **Batch** | `EvidenceCaseRunner` | Nightly eval, `run_question_bank.py` | ~17s (subprocess classifiers) | **Live mode**: YOU are the engine. Call `/memory recall` for data, do decomposition + entity analysis + same-technique check + verdict in your reasoning. No `/assistant classify` calls needed — you ARE the classifier. `/lean4-prove` Docker compilation is near-instant.
Ver en GitHub
Este SKILL.md es muy grande, por eso SkillsMP muestra aqui solo la primera seccion. Ver en GitHub