Skip to main content

extract-entities

Extract control IDs, domain phrases, taxonomy tags, and relationship data from question text. Composes /memory (ArangoDB recall to load sparta_controls vocabulary). Returns structured EntityExtractionResult that defines the shape of evidence cases, conversations, and QRA reviews before any LLM runs. Zero LLM cost — Flashtext (Aho-Corasick) + RapidFuzz (Levenshtein). NO REGEX.

소스 정보

저장소
grahama1970/agent-stack-public
최근 소스 활동
2026년 9월 24일 15:51
감지된 SKILL.md 언어
영어
스타
0
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
12 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
extract-entities
description
Extract control IDs, domain phrases, taxonomy tags, and relationship data from question text. Composes /memory (ArangoDB recall to load sparta_controls vocabulary). Returns structured EntityExtractionResult that defines the shape of evidence cases, conversations, and QRA reviews before any LLM runs. Zero LLM cost — Flashtext (Aho-Corasick) + RapidFuzz (Levenshtein). NO REGEX.
allowed-tools
["Bash","Read"]
triggers
["extract entities","what entities","what controls","decompose question","parse question"]
metadata
{"short-description":"Extract controls, phrases, and relationships from question text","author":"Horus","version":"0.1.0"}
provides
["entity-extraction"]
composes
["memory","taxonomy","analytics","agentic-evals"]
disciplines
["extraction","memory-knowledge"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT. # /extract-entities Extract control IDs, domain phrases, control metadata, relationship edges, interface commands, dataset/source mentions, and taxonomy tags from any question text. The extraction result defines the shape of the request before routing, gates, or LLM calls run. ## Always-On Request Gate Run `/extract-entities` on every user request in chat and project-agent systems. It is deterministic and cheap, so it should run before routing to `$ask`, `$create-figure`, `$create-evidence-case`, `$analytics`, or project-specific subagents. The result should be persisted with the run as `entity_context.json` so failures are debuggable. Downstream skills must read this structured context instead of regex-parsing the prompt again. For figure/data requests, extraction should identify: - skill and interface commands (`$create-figure`, `$analytics`, D3, graph, chart) - dataset names, Hugging Face dataset ids, splits/configs, file-like paths, and artifact references - requested variables, labels, controls, domain terms, unresolved terms, and explicit sample/demo permission For evidence requests, extraction should identify: - controls, standards, framework IDs, domain terms, relationships, and unresolved/fabricated-looking terms ## Architecture (NO REGEX, NO LLM) ### Why Flashtext Instead of ArangoDB Search? **ArangoDB text analyzers LEMMATIZE and TOKENIZE input.** "CWE-79" becomes two tokens: `cwe` and `79` (split on hyphen). This breaks exact entity matching. Flashtext does **exact string matching on raw text** before any tokenization. It finds "CWE-79" as a single unit. ArangoDB BM25/text search is used in Step 4 for finding **related content** after we already know what entities exist via Flashtext/RapidFuzz. ### The Flow The extraction flow is purely deterministic: ``` ┌─────────────────────────────────────────────────────────────────┐ │ Step 1: Load Vocabulary │ │ /memory recall → get ALL sparta_controls (~8,000 entries) │ │ → Load control_ids + names into Flashtext KeywordProcessor │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 2: Flashtext (Aho-Corasick) │ │ Run on RAW question text (no tokenization) │ │ → Returns exact matches with positions │ │ Example: "CWE-79" found at [32:38] │ │ These become PROTECTED TERMS │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 3: RapidFuzz (Levenshtein Distance) │ │ For unmatched terms that might be typos │ │ → "CWE-7" suggests "CWE-79" (distance: 1) │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 4: Hybrid Search (BM25 + Dense) │ │ Search question against sparta collections │ │ → sparta_qra, sparta_controls, sparta_url_knowledge │ │ → Returns recall_items (evidence) │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Step 5: spaCy Noun Phrase Extraction + Truncation │ │ Extract noun phrases from remaining text │ │ (excluding protected terms from Steps 2-3) │ │ → "ham sandwiches relate" extracted │ │ → Truncate glue words (NLTK POS tagging): │ │ - Strip trailing verbs: "relate" (VBP) removed │ │ - Strip leading articles/prepositions │ │ → "ham sandwiches" = cleaned phrase │ │ → Check against corpus (recall_items) │ │ → "ham sandwiches" = not in corpus = ungrounded │ │ │ │ OUTPUT: Same JSON structure as Flashtext entities: │ │ resolution_map["ham sandwiches"] = { │ │ exists: false, │ │ in_corpus: false, │ │ match_type: "noun_phrase" │ │ } │ └─────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────┐ │ Output: EntityExtractionResult │ │ control_ids: ["CWE-79"] │ │ resolution_map: { │ │ "CWE-79": {exists: true, match_type: "exact"}, │ │ "ham sandwiches": {exists: false, in_corpus: false} │ │ } │ │ recall_items: [QRAs about XSS...] │ └─────────────────────────────────────────────────────────────────┘ ``` **Critical rules:** - NO REGEX — Flashtext (Aho-Corasick), RapidFuzz (Levenshtein), spaCy (NLP) - NO LLM — extraction is 100% deterministic - NO manual tokenization — Flashtext runs on raw text, spaCy handles NLP - Vocabulary comes from ArangoDB via /memory, not hardcoded lists - Protected terms (from Flashtext/RapidFuzz) are excluded before spaCy runs - If ArangoDB misses an exact SPARTA control ID, control name, category, or alias, `/extract-entities` may perform a bounded lookup against the local authoritative `SPARTA-Data.xlsx` workbook. This is a stale-corpus fallback: returned entities must carry `source: "sparta_source_workbook_fallback"` or workbook provenance fields, and monitor/backfill should repair `$memory` so the normal ArangoDB path succeeds next time. ## Usage ```bash # Extract entities from a question ./run.sh extract "What does radar spoofing have to do with SV-AC-2 and CWE-89?" # Include taxonomy bridge attributes ./run.sh extract --taxonomy "How do NIST 800-171 requirements align with SPARTA defenses?" # JSON output for piping ./run.sh extract --json "Tell me about Control SV-CF-1 as it relates to D3FEND" # Compact project-agent output is the default JSON view ./run.sh extract --json "What is SPARTA countermeasure CM0029 (Comms Link)?" # Full diagnostic output for debugging old fields, glossary, and entity nodes ./run.sh extract --json --verbose "What is SPARTA countermeasure CM0029 (Comms Link)?" # Explicit compatibility mode for legacy consumers ./run.sh extract --json --view legacy "What is SPARTA countermeasure CM0029 (Comms Link)?" # Resolve entities from free text (NLP mode, default) ./run.sh resolve "radar spoofing impacts sensor fusion" # Resolve entities from delimited tokens (auto: comma/semicolon/whitespace) ./run.sh resolve --delimiter auto "SV-AC-2, CWE-89 radar_spoofing" # Resolve entities from custom delimiter-separated tokens ./run.sh resolve --delimiter "|" "SV-AC-2|CWE-89|unknown_token" ``` `resolve --delimiter` modes: - `nlp` (default): FlashText + fuzzy matching against collection dictionary - `auto`: split input by commas, semicolons, and whitespace, then lookup each token via `/list` - custom string: split on the provided delimiter, then lookup each token via `/list` ### Default stdin mode (no subcommand) When invoked without a subcommand, the script reads from stdin. Use `--delimiter` to control entity extraction mode and `--collection` to target any ArangoDB collection. ```bash # Delimiter mode: split tokens, lookup each by control_id filter echo "CA-7,PM-6,REC-0001" | python3 extract_entities.py \ --delimiter auto --collection sparta_controls # NLP mode (default when --delimiter is omitted): FlashText + fuzzy over collection echo "What countermeasures for supply chain?" | python3 extract_entities.py \ --collection sparta_controls ``` **Stdin flags:** | Flag | Default | Description | |------|---------|-------------| | `--delimiter` / `-d` | *(omit for NLP)* | `auto` splits on `,;` + whitespace; any other string is used as literal delimiter | | `--collection` / `-c` | `sparta_controls` | ArangoDB collection to match against | | `--name-field` | `name` | Field containing human-readable entity name | | `--label-field` | `control_id` | Field used as display label (e.g. `control_id`) | | `--framework-field` | `source_framework` | Field containing framework name | | `--type-field` | `node_type` | Field containing entity type | | `--limit` | `500` | Max entity docs loaded for NLP FlashText dictionary | | `--scope` | *(empty)* | Scope filter for `/recall` enrichment (NLP mode only) | **Delimiter mode lookup:** For each token, first issues `POST /list {"collection": …, "limit": 1, "filters": {"control_id": token}, "return_fields": ["control_id", "name", "source_framework"]}`. Falls back to `q`-based search with exact name/label match if the filter returns nothing. **Output shape (both modes):** ```json { "text": "CA-7,PM-6,REC-0001", "collection": "sparta_controls", "delimiter": "auto", "entity_count": 3, "entities": [ {"token": "CA-7", "id": "…", "name": "Continuous Monitoring", "label": "CA-7", "framework": "NIST", "exists": true}, {"token": "PM-6", "id": "…", "name": "…", "label": "PM-6", "framework": "NIST", "exists": true}, {"token": "REC-0001", "id": "", "name": "REC-0001", "label": "", "framework": "", "exists": false} ], "entity_names": ["Continuous Monitoring", "…", "REC-0001"], "entity_ids": ["sparta_controls/…", "sparta_controls/…"] } ``` ## Output ```json { "control_ids": ["SV-AC-2", "CWE-89"], "phrases": ["radar spoofing"], "phrase_controls": ["SV-CF-1", "SV-CF-3"], "all_control_ids": ["CWE-89", "SV-AC-2", "SV-CF-1", "SV-CF-3"], "control_metadata": [ {"control_id": "SV-AC-2", "name": "Access Control", "framework": "SPARTA", "domain": "..."} ], "glossary": [ { "id": "SV-AC-2", "control_id": "SV-AC-2", "name": "Access Control", "mention": "SV-AC-2", "framework": "SPARTA", "type": "countermeasure", "description": "...", "definition": "...", "aliases": [], "span": [42, 49], "grounded": true, "exists": true, "source": "sparta_controls" }, { "id": "descriptor:cm0029_comms_link", "mention": "Comms Link", "name": "Comms Link", "control_id": "CM0029", "type": "control_descriptor", "span": [38, 48], "grounded": true, "exists": true, "source": "sparta_controls_parenthetical_guard", "descriptor_kind": "control_category", "canonical_name": "TRANSEC", "category": "Comms Link" } ], "entity_nodes": [ { "id": "SV-AC-2", "node_kind": "control", "status": "grounded", "proof_role": "entity_anchor", "extracted": { "text": "SV-AC-2", "span": [42, 49], "source": "sparta_controls", "kind": "control_id" }, "metadata": { "control_id": "SV-AC-2", "name": "Access Control", "exists": true } }, { "id": "descriptor:cm0029_comms_link", "node_kind": "control_descriptor", "status": "grounded", "proof_role": "validated_context", "extracted": { "text": "Comms Link", "span": [38, 48], "source": "sparta_controls_parenthetical_guard", "kind": "control_descriptor" }, "metadata": { "name": "Comms Link", "exists": true, "control_id": "CM0029", "descriptor_kind": "control_category", "canonical_name": "TRANSEC", "category": "Comms Link" } }, { "id": "domain:satellite_uplink", "node_kind": "domain_term", "status": "extracted", "proof_role": "query_context", "extracted": { "text": "satellite uplink", "span": [32, 48], "source": "extract_entities_domain_terms", "kind": "domain_term" }, "metadata": { "name": "satellite uplink", "exists": true } } ], "related_pairs": [ {"source": "SV-AC-2", "target": "SV-CF-1", "method": "mitigates"} ], "taxonomy_tags": {"sparta": ["Signal_Manipulation"], "behavioral": ["Corruption"]}, "unresolved_terms": [ {"term": "X23-MUSTARD", "type": "id_like", "exists": false, "reason": "no_match_in_sparta_controls", "closest_match": "CM0028", "distance": 0.85} ], "resolution_map": { "SV-AC-2": {"exists": true, "match_type": "exact", "control_id": "SV-AC-2", "name": "Access Control", "qra_count": 14},
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기