Skip to main content

extract-entities

Extract control IDs, domain phrases, taxonomy tags, and relationship data from question text. Composes /memory (ArangoDB recall to load sparta_controls vocabulary). Returns structured EntityExtractionResult that defines the shape of evidence cases, conversations, and QRA reviews before any LLM runs. Zero LLM cost — Flashtext (Aho-Corasick) + RapidFuzz (Levenshtein). NO REGEX.

Ir para a instalação

Informações da origem

Repositório
grahama1970/agent-skills
Última atividade na origem
12 de agosto de 2026 às 15:01
Idioma detectado do SKILL.md
inglês
Estrelas
5
Forks
2

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Explorador de arquivos
13 arquivos

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
extract-entities
description
Extract control IDs, domain phrases, taxonomy tags, and relationship data from question text. Composes /memory (ArangoDB recall to load sparta_controls vocabulary). Returns structured EntityExtractionResult that defines the shape of evidence cases, conversations, and QRA reviews before any LLM runs. Zero LLM cost — Flashtext (Aho-Corasick) + RapidFuzz (Levenshtein). NO REGEX.
allowed-tools
["Bash","Read"]
triggers
["extract entities","what entities","what controls","decompose question","parse question"]
metadata
{"short-description":"Extract controls, phrases, and relationships from question text","author":"Horus","version":"0.1.0"}
provides
["entity-extraction"]
composes
["memory","taxonomy","analytics","agentic-evals"]
disciplines
["extraction","memory-knowledge"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT. # /extract-entities Extract control IDs, domain phrases, control metadata, relationship edges, interface commands, dataset/source mentions, and taxonomy tags from any question text. The extraction result defines the shape of the request before routing, gates, or LLM calls run. ## Always-On Request Gate Run `/extract-entities` on every user request in chat and project-agent systems. It is deterministic and cheap, so it should run before routing to `$ask`, `$create-figure`, `$create-evidence-case`, `$analytics`, or project-specific subagents. The result should be persisted with the run as `entity_context.json` so failures are debuggable. Downstream skills must read this structured context instead of regex-parsing the prompt again. For figure/data requests, extraction should identify: - skill and interface commands (`$create-figure`, `$analytics`, D3, graph, chart) - dataset names, Hugging Face dataset ids, splits/configs, file-like paths, and artifact references - requested variables, labels, controls, domain terms, unresolved terms, and explicit sample/demo permission For evidence requests, extraction should identify: - controls, standards, framework IDs, domain terms, relationships, and unresolved/fabricated-looking terms ## Architecture (NO REGEX, NO LLM) ### Why Flashtext Instead of ArangoDB Search? **ArangoDB text analyzers LEMMATIZE and TOKENIZE input.** "CWE-79" becomes two tokens: `cwe` and `79` (split on hyphen). This breaks exact entity matching. Flashtext does **exact string matching on raw text** before any tokenization. It finds "CWE-79" as a single unit. ArangoDB BM25/text search is used in Step 4 for finding **related content** after we already know what entities exist via Flashtext/RapidFuzz. ### The Flow The extraction flow is purely deterministic: ``` ┌─────────────────────────────────────────────────────────────────┐ │ Step 1: Load Vocabulary │ │ /memory recall → get ALL sparta_controls (~8,000 entries) │ │ → Load control_ids + names into Flashtext KeywordProcessor │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 2: Flashtext (Aho-Corasick) │ │ Run on RAW question text (no tokenization) │ │ → Returns exact matches with positions │ │ Example: "CWE-79" found at [32:38] │ │ These become PROTECTED TERMS │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 3: RapidFuzz (Levenshtein Distance) │ │ For unmatched terms that might be typos │ │ → "CWE-7" suggests "CWE-79" (distance: 1) │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 4: Hybrid Search (BM25 + Dense) │ │ Search question against sparta collections │ │ → sparta_qra, sparta_controls, sparta_url_knowledge │ │ → Returns recall_items (evidence) │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 5: spaCy Noun Phrase Extraction + Truncation │ │ Extract noun phrases from remaining text │ │ (excluding protected terms from Steps 2-3) │ │ → "ham sandwiches relate" extracted │ │ → Truncate glue words (NLTK POS tagging): │ │ - Strip trailing verbs: "relate" (VBP) removed │ │ - Strip leading articles/prepositions │ │ → "ham sandwiches" = cleaned phrase │ │ → Check against corpus (recall_items) │ │ → "ham sandwiches" = not in corpus = ungrounded │ │ │ │ OUTPUT: Same JSON structure as Flashtext entities: │ │ resolution_map["ham sandwiches"] = { │ │ exists: false, │ │ in_corpus: false, │ │ match_type: "noun_phrase" │ │ } │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Output: EntityExtractionResult │ │ control_ids: ["CWE-79"] │ │ resolution_map: { │ │ "CWE-79": {exists: true, match_type: "exact"}, │ │ "ham sandwiches": {exists: false, in_corpus: false} │ │ } │ │ recall_items: [QRAs about XSS...] │ └─────────────────────────────────────────────────────────────────┘ ``` **Critical rules:** - NO REGEX — Flashtext (Aho-Corasick), RapidFuzz (Levenshtein), spaCy (NLP) - NO LLM — extraction is 100% deterministic - NO manual tokenization — Flashtext runs on raw text, spaCy handles NLP - Vocabulary comes from ArangoDB via /memory, not hardcoded lists - Protected terms (from Flashtext/RapidFuzz) are excluded before spaCy runs - If ArangoDB misses an exact SPARTA control ID, control name, category, or alias, `/extract-entities` may perform a bounded lookup against the local authoritative `SPARTA-Data.xlsx` workbook. This is a stale-corpus fallback: returned entities must carry `source: "sparta_source_workbook_fallback"` or workbook provenance fields, and monitor/backfill should repair `$memory` so the normal ArangoDB path succeeds next time. ## Usage ```bash # Extract entities from a question ./run.sh extract "What does radar spoofing have to do with SV-AC-2 and CWE-89?" # Include taxonomy bridge attributes ./run.sh extract --taxonomy "How do NIST 800-171 requirements align with SPARTA defenses?" # JSON output for piping ./run.sh extract --json "Tell me about Control SV-CF-1 as it relates to D3FEND" # Compact project-agent output is the default JSON view ./run.sh extract --json "What is SPARTA countermeasure CM0029 (Comms Link)?" # Full diagnostic output for debugging old fields, glossary, and entity nodes ./run.sh extract --json --verbose "What is SPARTA countermeasure CM0029 (Comms Link)?" # Explicit compatibility mode for legacy consumers ./run.sh extract --json --view legacy "What is SPARTA countermeasure CM0029 (Comms Link)?" # Resolve entities from free text (NLP mode, default) ./run.sh resolve "radar spoofing impacts sensor fusion" # Resolve entities from delimited tokens (auto: comma/semicolon/whitespace) ./run.sh resolve --delimiter auto "SV-AC-2, CWE-89 radar_spoofing" # Resolve entities from custom delimiter-separated tokens ./run.sh resolve --delimiter "|" "SV-AC-2|CWE-89|unknown_token" ``` `resolve --delimiter` modes: - `nlp` (default): FlashText + fuzzy matching against collection dictionary - `auto`: split input by commas, semicolons, and whitespace, then lookup each token via `/list` - custom string: split on the provided delimiter, then lookup each token via `/list` ### Default stdin mode (no subcommand) When invoked without a subcommand, the script reads from stdin. Use `--delimiter` to control entity extraction mode and `--collection` to target any ArangoDB collection. ```bash # Delimiter mode: split tokens, lookup each by control_id filter echo "CA-7,PM-6,REC-0001" | python3 extract_entities.py \ --delimiter auto --collection sparta_controls # NLP mode (default when --delimiter is omitted): FlashText + fuzzy over collection echo "What countermeasures for supply chain?" | python3 extract_entities.py \ --collection sparta_controls ``` **Stdin flags:** | Flag | Default | Description | |------|---------|-------------| | `--delimiter` / `-d` | *(omit for NLP)* | `auto` splits on `,;` + whitespace; any other string is used as literal delimiter | | `--collection` / `-c` | `sparta_controls` | ArangoDB collection to match against | | `--name-field` | `name` | Field containing human-readable entity name | | `--label-field` | `control_id` | Field used as display label (e.g. `control_id`) | | `--framework-field` | `source_framework` | Field containing framework name | | `--type-field` | `node_type` | Field containing entity type | | `--limit` | `500` | Max entity docs loaded for NLP FlashText dictionary | | `--scope` | *(empty)* | Scope filter for `/recall` enrichment (NLP mode only) | **Delimiter mode lookup:** For each token, first issues `POST /list {"collection": …, "limit": 1, "filters": {"control_id": token}, "return_fields": ["control_id", "name", "source_framework"]}`. Falls back to `q`-based search with exact name/label match if the filter returns nothing. **Output shape (both modes):** ```json { "text": "CA-7,PM-6,REC-0001", "collection": "sparta_controls", "delimiter": "auto", "entity_count": 3, "entities": [ {"token": "CA-7", "id": "…", "name": "Continuous Monitoring", "label": "CA-7", "framework": "NIST", "exists": true}, {"token": "PM-6", "id": "…", "name": "…", "label": "PM-6", "framework": "NIST", "exists": true}, {"token": "REC-0001", "id": "", "name": "REC-0001", "label": "", "framework": "", "exists": false} ], "entity_names": ["Continuous Monitoring", "…", "REC-0001"], "entity_ids": ["sparta_controls/…", "sparta_controls/…"] } ``` ## Output ```json { "control_ids": ["SV-AC-2", "CWE-89"], "phrases": ["radar spoofing"], "phrase_controls": ["SV-CF-1", "SV-CF-3"], "all_control_ids": ["CWE-89", "SV-AC-2", "SV-CF-1", "SV-CF-3"], "control_metadata": [ {"control_id": "SV-AC-2", "name": "Access Control", "framework": "SPARTA", "domain": "..."} ], "glossary": [ { "id": "SV-AC-2", "control_id": "SV-AC-2", "name": "Access Control", "mention": "SV-AC-2", "framework": "SPARTA", "type": "countermeasure", "description": "...", "definition": "...", "aliases": [], "span": [42, 49], "grounded": true, "exists": true, "source": "sparta_controls" }, { "id": "descriptor:cm0029_comms_link", "mention": "Comms Link", "name": "Comms Link", "control_id": "CM0029", "type": "control_descriptor", "span": [38, 48], "grounded": true, "exists": true, "source": "sparta_controls_parenthetical_guard", "descriptor_kind": "control_category", "canonical_name": "TRANSEC", "category": "Comms Link" } ], "entity_nodes": [ { "id": "SV-AC-2", "node_kind": "control", "status": "grounded", "proof_role": "entity_anchor", "extracted": { "text": "SV-AC-2", "span": [42, 49], "source": "sparta_controls", "kind": "control_id" }, "metadata": { "control_id": "SV-AC-2", "name": "Access Control", "exists": true } }, { "id": "descriptor:cm0029_comms_link", "node_kind": "control_descriptor", "status": "grounded", "proof_role": "validated_context", "extracted": { "text": "Comms Link", "span": [38, 48], "source": "sparta_controls_parenthetical_guard", "kind": "control_descriptor" }, "metadata": { "name": "Comms Link", "exists": true, "control_id": "CM0029", "descriptor_kind": "control_category", "canonical_name": "TRANSEC", "category": "Comms Link" } }, { "id": "domain:satellite_uplink", "node_kind": "domain_term", "status": "extracted", "proof_role": "query_context", "extracted": { "text": "satellite uplink", "span": [32, 48], "source": "extract_entities_domain_terms", "kind": "domain_term" }, "metadata": { "name": "satellite uplink", "exists": true } } ], "related_pairs": [ {"source": "SV-AC-2", "target": "SV-CF-1", "method": "mitigates"} ], "taxonomy_tags": {"sparta": ["Signal_Manipulation"], "behavioral": ["Corruption"]}, "unresolved_terms": [ {"term": "X23-MUSTARD", "type": "id_like", "exists": false, "reason": "no_match_in_sparta_controls", "closest_match": "CM0028", "distance": 0.85} ], "resolution_map": { "SV-AC-2": {"exists": true, "match_type": "exact", "control_id": "SV-AC-2", "name": "Access Control", "qra_count": 14},
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub