- name
- extract-entities
- description
- Extract control IDs, domain phrases, taxonomy tags, and relationship data from question text. Composes /memory (ArangoDB recall to load sparta_controls vocabulary). Returns structured EntityExtractionResult that defines the shape of evidence cases, conversations, and QRA reviews before any LLM runs. Zero LLM cost — Flashtext (Aho-Corasick) + RapidFuzz (Levenshtein). NO REGEX.
- allowed-tools
- ["Bash","Read"]
- triggers
- ["extract entities","what entities","what controls","decompose question","parse question"]
- metadata
- {"short-description":"Extract controls, phrases, and relationships from question text","author":"Horus","version":"0.1.0"}
- provides
- ["entity-extraction"]
- composes
- ["memory","taxonomy","analytics","agentic-evals"]
- disciplines
- ["extraction","memory-knowledge"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.
# /extract-entities
Extract control IDs, domain phrases, control metadata, relationship edges, interface commands, dataset/source mentions, and taxonomy tags from any question text. The extraction result defines the shape of the request before routing, gates, or LLM calls run.
## Always-On Request Gate
Run `/extract-entities` on every user request in chat and project-agent systems.
It is deterministic and cheap, so it should run before routing to `$ask`,
`$create-figure`, `$create-evidence-case`, `$analytics`, or project-specific
subagents.
The result should be persisted with the run as `entity_context.json` so failures
are debuggable. Downstream skills must read this structured context instead of
regex-parsing the prompt again.
For figure/data requests, extraction should identify:
- skill and interface commands (`$create-figure`, `$analytics`, D3, graph, chart)
- dataset names, Hugging Face dataset ids, splits/configs, file-like paths, and artifact references
- requested variables, labels, controls, domain terms, unresolved terms, and explicit sample/demo permission
For evidence requests, extraction should identify:
- controls, standards, framework IDs, domain terms, relationships, and unresolved/fabricated-looking terms
## Architecture (NO REGEX, NO LLM)
### Why Flashtext Instead of ArangoDB Search?
**ArangoDB text analyzers LEMMATIZE and TOKENIZE input.** "CWE-79" becomes two tokens: `cwe` and `79` (split on hyphen). This breaks exact entity matching.
Flashtext does **exact string matching on raw text** before any tokenization. It finds "CWE-79" as a single unit.
ArangoDB BM25/text search is used in Step 4 for finding **related content** after we already know what entities exist via Flashtext/RapidFuzz.
### The Flow
The extraction flow is purely deterministic:
```
┌─────────────────────────────────────────────────────────────────┐
│ Step 1: Load Vocabulary │
│ /memory recall → get ALL sparta_controls (~8,000 entries) │
│ → Load control_ids + names into Flashtext KeywordProcessor │
└─────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Step 2: Flashtext (Aho-Corasick) │
│ Run on RAW question text (no tokenization) │
│ → Returns exact matches with positions │
│ Example: "CWE-79" found at [32:38] │
│ These become PROTECTED TERMS │
└─────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Step 3: RapidFuzz (Levenshtein Distance) │
│ For unmatched terms that might be typos │
│ → "CWE-7" suggests "CWE-79" (distance: 1) │
└─────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Step 4: Hybrid Search (BM25 + Dense) │
│ Search question against sparta collections │
│ → sparta_qra, sparta_controls, sparta_url_knowledge │
│ → Returns recall_items (evidence) │
└─────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Step 5: spaCy Noun Phrase Extraction + Truncation │
│ Extract noun phrases from remaining text │
│ (excluding protected terms from Steps 2-3) │
│ → "ham sandwiches relate" extracted │
│ → Truncate glue words (NLTK POS tagging): │
│ - Strip trailing verbs: "relate" (VBP) removed │
│ - Strip leading articles/prepositions │
│ → "ham sandwiches" = cleaned phrase │
│ → Check against corpus (recall_items) │
│ → "ham sandwiches" = not in corpus = ungrounded │
│ │
│ OUTPUT: Same JSON structure as Flashtext entities: │
│ resolution_map["ham sandwiches"] = { │
│ exists: false, │
│ in_corpus: false, │
│ match_type: "noun_phrase" │
│ } │
└─────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Output: EntityExtractionResult │
│ control_ids: ["CWE-79"] │
│ resolution_map: { │
│ "CWE-79": {exists: true, match_type: "exact"}, │
│ "ham sandwiches": {exists: false, in_corpus: false} │
│ } │
│ recall_items: [QRAs about XSS...] │
└─────────────────────────────────────────────────────────────────┘
```
**Critical rules:**
- NO REGEX — Flashtext (Aho-Corasick), RapidFuzz (Levenshtein), spaCy (NLP)
- NO LLM — extraction is 100% deterministic
- NO manual tokenization — Flashtext runs on raw text, spaCy handles NLP
- Vocabulary comes from ArangoDB via /memory, not hardcoded lists
- Protected terms (from Flashtext/RapidFuzz) are excluded before spaCy runs
- If ArangoDB misses an exact SPARTA control ID, control name, category, or
alias, `/extract-entities` may perform a bounded lookup against the local
authoritative `SPARTA-Data.xlsx` workbook. This is a stale-corpus fallback:
returned entities must carry `source: "sparta_source_workbook_fallback"` or
workbook provenance fields, and monitor/backfill should repair `$memory` so
the normal ArangoDB path succeeds next time.
## Usage
```bash
# Extract entities from a question
./run.sh extract "What does radar spoofing have to do with SV-AC-2 and CWE-89?"
# Include taxonomy bridge attributes
./run.sh extract --taxonomy "How do NIST 800-171 requirements align with SPARTA defenses?"
# JSON output for piping
./run.sh extract --json "Tell me about Control SV-CF-1 as it relates to D3FEND"
# Compact project-agent output is the default JSON view
./run.sh extract --json "What is SPARTA countermeasure CM0029 (Comms Link)?"
# Full diagnostic output for debugging old fields, glossary, and entity nodes
./run.sh extract --json --verbose "What is SPARTA countermeasure CM0029 (Comms Link)?"
# Explicit compatibility mode for legacy consumers
./run.sh extract --json --view legacy "What is SPARTA countermeasure CM0029 (Comms Link)?"
# Resolve entities from free text (NLP mode, default)
./run.sh resolve "radar spoofing impacts sensor fusion"
# Resolve entities from delimited tokens (auto: comma/semicolon/whitespace)
./run.sh resolve --delimiter auto "SV-AC-2, CWE-89 radar_spoofing"
# Resolve entities from custom delimiter-separated tokens
./run.sh resolve --delimiter "|" "SV-AC-2|CWE-89|unknown_token"
```
`resolve --delimiter` modes:
- `nlp` (default): FlashText + fuzzy matching against collection dictionary
- `auto`: split input by commas, semicolons, and whitespace, then lookup each token via `/list`
- custom string: split on the provided delimiter, then lookup each token via `/list`
### Default stdin mode (no subcommand)
When invoked without a subcommand, the script reads from stdin. Use `--delimiter` to
control entity extraction mode and `--collection` to target any ArangoDB collection.
```bash
# Delimiter mode: split tokens, lookup each by control_id filter
echo "CA-7,PM-6,REC-0001" | python3 extract_entities.py \
--delimiter auto --collection sparta_controls
# NLP mode (default when --delimiter is omitted): FlashText + fuzzy over collection
echo "What countermeasures for supply chain?" | python3 extract_entities.py \
--collection sparta_controls
```
**Stdin flags:**
| Flag | Default | Description |
|------|---------|-------------|
| `--delimiter` / `-d` | *(omit for NLP)* | `auto` splits on `,;` + whitespace; any other string is used as literal delimiter |
| `--collection` / `-c` | `sparta_controls` | ArangoDB collection to match against |
| `--name-field` | `name` | Field containing human-readable entity name |
| `--label-field` | `control_id` | Field used as display label (e.g. `control_id`) |
| `--framework-field` | `source_framework` | Field containing framework name |
| `--type-field` | `node_type` | Field containing entity type |
| `--limit` | `500` | Max entity docs loaded for NLP FlashText dictionary |
| `--scope` | *(empty)* | Scope filter for `/recall` enrichment (NLP mode only) |
**Delimiter mode lookup:** For each token, first issues
`POST /list {"collection": …, "limit": 1, "filters": {"control_id": token}, "return_fields": ["control_id", "name", "source_framework"]}`.
Falls back to `q`-based search with exact name/label match if the filter returns nothing.
**Output shape (both modes):**
```json
{
"text": "CA-7,PM-6,REC-0001",
"collection": "sparta_controls",
"delimiter": "auto",
"entity_count": 3,
"entities": [
{"token": "CA-7", "id": "…", "name": "Continuous Monitoring", "label": "CA-7", "framework": "NIST", "exists": true},
{"token": "PM-6", "id": "…", "name": "…", "label": "PM-6", "framework": "NIST", "exists": true},
{"token": "REC-0001", "id": "", "name": "REC-0001", "label": "", "framework": "", "exists": false}
],
"entity_names": ["Continuous Monitoring", "…", "REC-0001"],
"entity_ids": ["sparta_controls/…", "sparta_controls/…"]
}
```
## Output
```json
{
"control_ids": ["SV-AC-2", "CWE-89"],
"phrases": ["radar spoofing"],
"phrase_controls": ["SV-CF-1", "SV-CF-3"],
"all_control_ids": ["CWE-89", "SV-AC-2", "SV-CF-1", "SV-CF-3"],
"control_metadata": [
{"control_id": "SV-AC-2", "name": "Access Control", "framework": "SPARTA", "domain": "..."}
],
"glossary": [
{
"id": "SV-AC-2",
"control_id": "SV-AC-2",
"name": "Access Control",
"mention": "SV-AC-2",
"framework": "SPARTA",
"type": "countermeasure",
"description": "...",
"definition": "...",
"aliases": [],
"span": [42, 49],
"grounded": true,
"exists": true,
"source": "sparta_controls"
},
{
"id": "descriptor:cm0029_comms_link",
"mention": "Comms Link",
"name": "Comms Link",
"control_id": "CM0029",
"type": "control_descriptor",
"span": [38, 48],
"grounded": true,
"exists": true,
"source": "sparta_controls_parenthetical_guard",
"descriptor_kind": "control_category",
"canonical_name": "TRANSEC",
"category": "Comms Link"
}
],
"entity_nodes": [
{
"id": "SV-AC-2",
"node_kind": "control",
"status": "grounded",
"proof_role": "entity_anchor",
"extracted": {
"text": "SV-AC-2",
"span": [42, 49],
"source": "sparta_controls",
"kind": "control_id"
},
"metadata": {
"control_id": "SV-AC-2",
"name": "Access Control",
"exists": true
}
},
{
"id": "descriptor:cm0029_comms_link",
"node_kind": "control_descriptor",
"status": "grounded",
"proof_role": "validated_context",
"extracted": {
"text": "Comms Link",
"span": [38, 48],
"source": "sparta_controls_parenthetical_guard",
"kind": "control_descriptor"
},
"metadata": {
"name": "Comms Link",
"exists": true,
"control_id": "CM0029",
"descriptor_kind": "control_category",
"canonical_name": "TRANSEC",
"category": "Comms Link"
}
},
{
"id": "domain:satellite_uplink",
"node_kind": "domain_term",
"status": "extracted",
"proof_role": "query_context",
"extracted": {
"text": "satellite uplink",
"span": [32, 48],
"source": "extract_entities_domain_terms",
"kind": "domain_term"
},
"metadata": {
"name": "satellite uplink",
"exists": true
}
}
],
"related_pairs": [
{"source": "SV-AC-2", "target": "SV-CF-1", "method": "mitigates"}
],
"taxonomy_tags": {"sparta": ["Signal_Manipulation"], "behavioral": ["Corruption"]},
"unresolved_terms": [
{"term": "X23-MUSTARD", "type": "id_like", "exists": false,
"reason": "no_match_in_sparta_controls", "closest_match": "CM0028", "distance": 0.85}
],
"resolution_map": {
"SV-AC-2": {"exists": true, "match_type": "exact", "control_id": "SV-AC-2",
"name": "Access Control", "qra_count": 14},
Ver no GitHub