- name
- create-evidence-case
- description
- Build structured evidence cases using Claims-Arguments-Evidence (CAE) trees. AGENT-DRIVEN composable orchestrator: the agent decomposes the question, calls existing skills for data collection and verification, then DECIDES the verdict. Python runner.py is a thin data collector and persistence layer.
- allowed-tools
- ["Bash","Read","Grep","Glob","Write","Edit"]
- triggers
- ["create-evidence-case","evidence case","build evidence","verify claim","evidence tree"]
- metadata
- {"short-description":"Composable CAE evidence orchestrator","version":"4.3.0"}
- provides
- ["evidence-case-creation","uct-strategy-selection","evidence-tree-persistence","claim-verdict-grading","assurance-case-output","oscal-export","qra-candidate-quarantine","human-review-interview"]
- composes
- ["memory","extract-entities","assistant","taxonomy","lean4-prove","dogpile","edge-verifier","cmmc-assessor","create-gsn-diagram","export-oscal","create-figure","task-monitor","interview","prompt-lab","analytics","agentic-evals"]
- taxonomy
- ["verification","evidence","quality-assurance"]
- disciplines
- ["compliance-security","evaluation-quality"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.
**User-facing documentation:** See [COMPLETE_EXAMPLE.md](COMPLETE_EXAMPLE.md) for compliance officer workflow, CAE tree examples, and the tool's role as a research assistant (not compliance oracle).
# /create-evidence-case v4.3 — Composable Orchestrator
Build structured **Claims-Arguments-Evidence (CAE)** trees. This is an **agent-driven** skill — you (the agent) orchestrate existing skills, reason about results, and make all judgment calls. The Python code is a thin data collector.
## PROVIDER BOUNDARY
Do not call `$scillm`, `/scillm`, `http://localhost:4001`,
`/v1/chat/completions`, or `/v1/scillm/*` directly from a
`/create-evidence-case` project-agent workflow. Provider/model execution is
owned by Tau. If model rendering, batch review, shadow collection, or provider
delegation is required, express it as a Tau DAG contract, Tau skill node, or
the owning skill's Tau-mediated runtime and consume the resulting receipts.
Direct SciLLM use is allowed only for Tau/SciLLM maintenance or when the human
explicitly asks to operate SciLLM itself. A CAE verdict must remain grounded in
Memory recall, `/extract-entities` proof packets, same-technique checks, and
deterministic gates; raw model output is never evidence.
## INPUT/OUTPUT MODEL (v4.2)
**Input**: Any control from `sparta_controls` (SPARTA, NIST, CWE, CAPEC) OR a raw question.
**Output**: Always a QRA document in `sparta_qra`, regardless of input source.
| Input Source | Example | Output |
|--------------|---------|--------|
| SPARTA control | SA-1 | QRA with `source_framework: "SPARTA"` |
| NIST control | IA-5 | QRA with `source_framework: "NIST-800-53"` |
| CWE | CWE-287 | QRA with `source_framework: "CWE"` |
| Raw question | "What is CWE-287?" | QRA with extracted `source_control_id` |
### QRA Output Tiers
Not every QRA has evidence. The `evidence_case` field being `null` is a **deterministic signal** (no crosswalk edges exist), not a failure.
| Tier | evidence_case | Use Case |
|------|---------------|----------|
| **Informational** | null | Explanation, lookup (no crosswalk edges) |
| **Grounded** | `{chains: [...], formal_proof: null}` | Compliance, audit |
| **Verified** | `{chains: [...], formal_proof: {success: true}}` | Safety-critical, formal methods |
All fields (chains, formal_proof, sacm_ref) live **inside** evidence_case. One null check.
### Entity Extraction → Evidence Case Flow
```
"What is CWE-287?"
│
▼
/extract-entities → "CWE-287" ✓
│
▼
/create-evidence-case
├── Query crosswalk edges in sparta_relationships
├── Has edges → evidence_case: {chains: [...]}
└── No edges → evidence_case: null (valid output)
│
▼
QRA created in sparta_qra
├── question: "What is CWE-287?"
├── requirement: CWE-287 description
├── answer: Explanation text
├── source_framework: "CWE"
├── source_control_id: "CWE-287"
└── evidence_case: {...} or null # SELF-CONTAINED: includes chains, proof, sacm_ref
```
### Self-Contained QRA Schema (v4.3)
Every QRA embeds all metadata for downstream consumption. **evidence_case is fully
self-contained** — formal proofs and SACM references live INSIDE it, not at top level.
This ensures traceability in ArangoDB and portability for OSCAL/SACM export.
```python
{
# Core QRA fields
"question": str,
"requirement": str,
"answer": str,
"source_framework": str, # SPARTA, NIST-800-53, CWE, CAPEC
"source_control_id": str, # SA-1, IA-5, CWE-287
"relationship_id": str, # Links persona variations
"expertise": str, # expert, layperson, pm, compliance
"difficulty": str, # single_hop, multi_hop, synthesis
# Taxonomy (Mind tags for SPARTA scope)
"mind": list[str], # ["Detect", "Harden"]
# Evidence case — SELF-CONTAINED (null if no crosswalk edges)
# All verification evidence lives here: chains, proof, SACM ref
"evidence_case": {
"chains": [...], # Crosswalk paths (source → hops → target)
"confidence": float, # 0.0-1.0 combined confidence
"methods": list[str], # ["direct", "mitre_chain", "nist_nvd"]
# Formal proof (null if not verified) — INSIDE evidence_case
"formal_proof": {
"success": bool,
"code": str, # Lean4 source code
"attempts": int,
"errors": list | None,
"proved_at": int | None,
} | None,
# SACM reference (null if not exported) — INSIDE evidence_case
"sacm_ref": {
"gid": str, # SACM Goal node GID (lookup key)
"xml_snippet": str, # First 500 chars (preview)
"generated_at": int,
} | None,
} | None,
}
```
**Why self-contained?** ISO 26262 and DO-178C refinement checking patterns bundle
verification evidence with the claim. Separate collections would require joins,
lose portability, and break single-document export to OSCAL/SACM XML.
### Cached Evidence Case in QRAs (v4.3.1)
Every QRA has a pre-computed `evidence_case` field containing extracted entities
with **character spans** for UI inline highlighting. This is populated by the
`evidence_case_backfill` pipeline (not query-time extraction).
```python
{
"evidence_case": {
# Core extracted data (from /extract-entities)
"question_text": str, # First 500 chars of question
"control_ids": list[str], # ["IA-0006", "CWE-287"]
"glossary": [ # Snapshot copied from /extract-entities.glossary
{
"id": "IA-0006",
"name": "Authentication Mechanisms",
"framework": "SPARTA",
"type": "countermeasure",
"description": "..."
}
],
"crosswalk_chains": [...], # Framework-to-framework links
"resolved_entities": [...], # Terms that resolved to controls
# --- SPANS FOR UI HIGHLIGHTING (SPARTA Explorer Chat) ---
"spans": [
{
"text": "IA-0006", # The matched text
"span": [9, 16], # [start, end] character positions
"kind": "control_id", # Type: control_id, aerospace_term, phrase
"framework": "SPARTA", # Only for control_id kind
"name": "Authentication Mechanisms",
"grounded_to_framework": True
},
{
"text": "CWE-287",
"span": [33, 40],
"kind": "control_id",
"framework": "CWE",
"name": "Improper Authentication",
"grounded_to_framework": True
}
],
"review_status": "auto", # auto, approved, rejected
"extracted_at": str # ISO timestamp
}
}
```
### SPARTA Explorer Chat: Inline Evidence Highlighting
The Chat agent uses `evidence_case.spans` to render inline highlights on question text:
```tsx
// React component example
function QuestionWithHighlights({ question, spans }) {
// Sort spans by start position
const sorted = [...spans].sort((a, b) => a.span[0] - b.span[0]);
let result = [];
let lastEnd = 0;
for (const s of sorted) {
const [start, end] = s.span;
// Text before this span
if (start > lastEnd) {
result.push(<span key={`t-${lastEnd}`}>{question.slice(lastEnd, start)}</span>);
}
// The highlighted entity
const color = s.kind === 'control_id'
? FRAMEWORK_COLORS[s.framework]
: 'yellow';
result.push(
<Tooltip key={`e-${start}`} content={`${s.name} (${s.framework})`}>
<span className={`bg-${color}-200`}>{s.text}</span>
</Tooltip>
);
lastEnd = end;
}
// Remaining text
if (lastEnd < question.length) {
result.push(<span key="tail">{question.slice(lastEnd)}</span>);
}
return <p>{result}</p>;
}
```
**Span kinds and colors:**
| Kind | Color | Example |
|------|-------|---------|
| `control_id` | Framework-specific (SPARTA=blue, CWE=orange, NIST=green) | IA-0006, CWE-287 |
| `aerospace_term` | Purple | FPGA, TT&C, crosslink |
| `phrase` | Yellow | supply chain attack |
**Framework colors (NVIS standard):**
| Framework | Color | Hex |
|-----------|-------|-----|
| SPARTA | Blue | `#3B82F6` |
| CWE | Orange | `#F97316` |
| NIST | Green | `#22C55E` |
| CAPEC | Red | `#EF4444` |
| ATT&CK | Purple | `#A855F7` |
### Cached vs Query-Time Evidence
| Field | When Populated | Use |
|-------|----------------|-----|
| `evidence_case.spans` | Backfill (pre-computed) | UI highlighting |
| `evidence_case.glossary` | Backfill | Cached `/extract-entities.glossary` metadata |
| `cae_tree` (in endpoint response) | Query-time | Full CAE tree with traceability |
| `entity_context` (in endpoint response) | Query-time | Compact `/extract-entities` agent view with authoritative `proof_packet.assertions`, `nodes`, and candidate-only boundaries |
The cached `evidence_case` enables fast rendering without query-time entity extraction.
Glossary semantics are owned by `/extract-entities`; `/create-evidence-case`
copies that glossary into the evidence case and must not independently infer
control names, definitions, or unsupported parenthetical terms.
Runtime extraction proof is preserved in `entity_context`; agents should use
`entity_context.proof_packet.assertions` for grounding decisions and keep
legacy glossary fields for compatibility, prompts, and UI display.
The `cae_tree` (requested via `include_cae_tree=True`) provides full CAE analysis
with traceability strength and validation gates.
## TRACEABILITY vs VERIFICATION
Matching requirements to controls (via `/match/requirement`) provides **traceability** —
linking related concepts. But traceability alone is NOT SUFFICIENT for certification.
| Level | What it proves | Provides | Sufficient for |
|-------|----------------|----------|----------------|
| **Matching** | "Related concepts" | Traceability | Gap detection, triage |
| **Confidence scoring** | "Probably covers" | Prioritization | Audit prep |
| **Formal proof** | "Functionally entails" | Verification | ISO 26262, DO-178C certification |
**Why matching is insufficient:**
```
Requirement: "shall implement multi-factor authentication for all users"
Control: SV-AC-2 "implement access control mechanisms"
Matching confidence: 0.82
- Crosswalk chain exists (CWE → NIST → SPARTA)
- QRA evidence mentions both
- Semantic similarity is high
Problem: Does SV-AC-2 ACTUALLY cover MFA?
- Words overlap, but FUNCTIONS might not
- Control might be narrower/broader than requirement
- Edge cases might not be addressed
```
**Why formal proofs are needed:**
```
Formal proof (Lean4):
theorem control_satisfies_requirement :
∀ u : User, implements(SV-AC-2, u) → authenticated(u) ∧ factors(u) ≥ 2
Proof result: success=true
Meaning: The control ENTAILS the requirement — not just similar words,
but the control is a valid implementation of the requirement.
```
The `evidence_case.formal_proof` field stores this verification evidence.
Matching builds the edges; formal proofs prove they're functionally valid.
## EXECUTION MODES
| Mode | Engine | When | Latency |
|------|--------|------|---------|
| **Live** | Project agent (you) | Interactive use | ~10-12s (dominated by `/memory recall`) |
| **Batch** | `EvidenceCaseRunner` | Nightly eval, `run_question_bank.py` | ~17s (subprocess classifiers) |
**Live mode**: YOU are the engine. Call `/memory recall` for data, do decomposition + entity analysis + same-technique check + verdict in your reasoning. No `/assistant classify` calls needed — you ARE the classifier. `/lean4-prove` Docker compilation is near-instant.
عرض على GitHub