Skip to main content

ingest-code

Ingest codebases into /memory for knowledge extraction and CWE scanning. Phase 1 extracts functional knowledge (module docstrings, function signatures, class hierarchies, markdown docs) via Python AST. Phase 2 scans for CWE mappings via /taxonomy. Designed to run nightly via /monitor-codebase.

Aller à l'installation

Informations de source

Dépôt
grahama1970/agent-skills
Dernière activité de la source
11 août 2026 à 15:34
Langue détectée de SKILL.md
anglais
Étoiles
5
Forks
2

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
49 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
ingest-code
description
Ingest codebases into /memory for knowledge extraction and CWE scanning. Phase 1 extracts functional knowledge (module docstrings, function signatures, class hierarchies, markdown docs) via Python AST. Phase 2 scans for CWE mappings via /taxonomy. Designed to run nightly via /monitor-codebase.
allowed-tools
["Bash","Read"]
triggers
["ingest code","scan codebase","ingest codebase","scan for cwes","codebase ingestion","code to memory"]
metadata
{"short-description":"Codebase knowledge + CWE extraction to /memory","author":"Horus","version":"0.2.0"}
provides
["ingest-code"]
composes
["memory","taxonomy","task-monitor","agentic-evals"]
disciplines
["data-engineering","memory-knowledge"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT. # ingest-code Codebase ingestion into `/memory`: 1. **Phase 1 — Functional Knowledge**: Python AST extracts module docstrings, function signatures, class hierarchies. Markdown parser extracts section-level knowledge from CONTEXT.md, README.md, etc. Generic parser handles TS/JS exports. 2. **Phase 2 — CWE Scanning**: `/taxonomy` extracts security-relevant patterns (bridge tags + CWE mappings) per file. 3. **Phase 3 — Relationship Edges**: Python import analysis stores resolved code dependency edges in `/memory`; the local code-graph bundle also records typed import/call/inheritance occurrences, including inactive unresolved and ambiguous candidates for audit. 4. **Phase 4 — Structured Code Index**: Optional Tree-sitter extraction emits a complete code-graph bundle and applies it through Memory/GMO's governed `/code/projection/apply` lifecycle endpoint. 5. **Static Debugger Affordances**: Code-graph bundles emit `debugger.invocation_candidate.v1` rows in `debug_invocations.jsonl`. These are candidate routes only; `$debugger` must prove a candidate at runtime before Memory can promote it. Functional lessons and CWE summaries remain lesson-style memory records for compatibility. Structured files, symbols, and edges are canonicalized by Memory/GMO's complete projection lifecycle, not by independent per-symbol batches. `/memory` owns ArangoDB, Qdrant, embeddings, sparse/hybrid retrieval, and payload/index behavior. `/ingest-code` must not talk to ArangoDB or Qdrant directly. ## Quick Start ```bash cd .pi/skills/ingest-code # Full knowledge + CWE scan ./run.sh scan /path/to/codebase # CWE scan only (legacy mode) ./run.sh scan /path/to/codebase --cwe-only # Preview without storing to /memory ./run.sh scan /path/to/codebase --dry-run # Include Tree-sitter structured code symbols for memory's hybrid code index ./run.sh scan /path/to/codebase --treesitter # Nightly rescan (only files modified in last day) ./run.sh rescan --since 1d -c /path/to/codebase --treesitter # Read-only target freshness check before using Memory code snippets for repair ./run.sh ensure-current \ --repo /path/to/codebase \ --branch main \ --commit "$COMMIT_SHA" \ --path src/example.py \ --json ``` ## Commands ### `scan` — Full Codebase Scan ```bash ./run.sh scan <path> [OPTIONS] Options: --glob, -g File patterns (default: *.py *.ts *.js *.rs *.go *.java *.c *.cpp) --cwe-only Skip Phase 1, only run CWE scan --validate Run LLM validation on CWE matches --treesitter Run Tree-sitter scan for structured code symbols --code-index Apply the complete Tree-sitter code graph bundle to Memory/GMO (default) --no-code-index Disable structured code projection application --compat-symbol-upsert Use legacy per-symbol Memory upserts with a visible warning --dry-run Preview without writing to /memory --scope Memory scope (default: "code") --batch-size Files per CWE scan batch (default: 50) ``` ### `rescan` — Incremental Rescan (Scheduler Job) ```bash ./run.sh rescan [OPTIONS] Options: --since Only files modified since (ISO date or "1d", "7d") -c, --codebase Codebase path(s) to rescan (repeatable) --validate Run LLM validation --treesitter Run Tree-sitter scan for structured code symbols --code-index Apply the complete Tree-sitter code graph bundle to Memory/GMO (default) --no-code-index Disable structured code projection application --scope Memory scope ``` ### `ensure-current` — Target-Scoped Projection Freshness Preflight ```bash ./run.sh ensure-current [OPTIONS] Options: --repo Repository worktree to check (required) --branch Expected branch/ref name (default: current branch) --commit Expected commit SHA (default: current HEAD) --path Repository-relative target path; repeatable --scope Memory/GMO projection scope (default: "code") --json Emit `ingest-code.code_projection_freshness.v1` --refresh Explicitly refresh through `scan --treesitter --code-index` --canonical-branch Branch allowed to activate canonical projection (default: "main") --max-target-files Bound directory expansion (default: 200) ``` `ensure-current` is the pre-repair gate for stateless workers. It resolves the repository root, branch, commit, and target paths; rejects absolute paths, `..`, and repository escapes; reads active code-search/code-node/code-coverage state through the supported Memory/GMO code-navigation boundary; then compares current source hashes with indexed source hashes for the requested targets. The result status is one of: | Status | Meaning | |--------|---------| | `CURRENT` | Active Memory/GMO source hashes match current target files and coverage allows modification guidance. | | `SOURCE_CURRENT_INDEX_INCOMPLETE` | Source bytes match, but coverage is incomplete, so exhaustive callers/callees/impact absence claims are blocked. | | `STALE` | Current source differs from the active projection; stored snippets are not modification authority. | | `UNINDEXED` | No applicable active projection record matched the target. | | `BLOCKED` | Identity, containment, service, receipt, or validation failed closed. | Default `ensure-current` is read-only. It must not parse files, create embeddings, apply projections, or fall back to legacy per-symbol writes. With `--refresh`, it may run the existing complete-bundle scan only when the checkout is clean, on the configured canonical branch, and bound to the requested commit. Feature/repair worktrees are refused so unreviewed code cannot replace the canonical main projection. ## What Gets Extracted ### Phase 1: Functional Knowledge (Python AST) | Source | What | Example /memory Problem | |--------|------|------------------------| | Module docstring | Module purpose | "What does run_pipeline.py do?" | | Class definition | Class + methods + bases | "What is the ContentRepository class in content_query.py?" | | Function signature | Args, return type, docstring | "What does extract_tables() do in s05_table_extractor.py?" | | Markdown sections | Architecture decisions, bug fixes | "What does 'Bugs Fixed' say in MEMORY.md?" | | TS/JS exports | Exported symbols | "What is AnswerCanvas in AnswerCanvas.tsx?" | ### Phase 2: CWE Scanning (via /taxonomy) | Category | Example CWEs | Triggers | |----------|--------------|----------| | MemorySafety | CWE-120, CWE-787, CWE-416 | buffer, overflow, memory, pointer | | InputValidation | CWE-20, CWE-89, CWE-78 | input, validation, inject, command | | Authentication | CWE-287, CWE-798, CWE-522 | auth, credential, password, session | | Cryptography | CWE-311, CWE-327, CWE-330 | encrypt, crypto, key, random | ### Phase 4: Structured Code Index (Tree-sitter → /memory) When `--treesitter --code-index` is enabled, `/ingest-code` writes a deterministic local code-graph bundle under `artifacts/ingest-code/code-graph/`, computes the submitted bundle digest and checksums digest, and submits the complete bundle to Memory/GMO through `/code/projection/apply`. The Memory/GMO receipt is the only canonical projection success signal. It must bind the submitted bundle digest, checksums digest, activated generation, and expected file/symbol/edge counts. If Memory/GMO is unavailable, rejects the bundle, or returns a receipt whose digest does not match the submitted bundle, the scan fails closed and does not fall back to per-symbol writes. Each record includes: | Field | Purpose | |-------|---------| | `repo`, `branch`, `commit`, `path` | Scope and staleness control | | `language`, `symbol_kind`, `symbol_name`, `qualified_name` | Symbol filtering and exact lookup | | `start_line`, `end_line`, `code`, `content_hash` | Cited source retrieval and deterministic updates | | `imports`, `parameters`, `local_variables`, `called_symbols`, `string_literals` | Lexical terms for memory's sparse/hybrid retrieval | | `problem`, `solution`, `text`, `tags` | Compatibility with existing memory recall surfaces | Documentation metadata is provenance-safe: | Field | Purpose | |-------|---------| | `source_docstring` / `docstring` | Exact authored source docstring text, preserved for compatibility | | `source_docstring_status` | `present`, `missing`, `generated_file`, or `not_applicable` for v1 extraction | | `documentation_need` | Deterministic triage: `required`, `recommended`, `optional`, or `exempt` | | `documentation_need_reasons` | Source-derived reasons such as `public_api`, `external_io`, `security`, `mutation`, or `trivial_helper` | | `summary_evidence` | Canonical source-fact packet and hash for optional generated summaries | | `derived_summary` | Current unreviewed generated summary only when bound to the current `symbol_version_id`, source hash, and evidence hash | | `retrieval_text`, `retrieval_text_sha256`, `purpose_source` | Canonical semantic text and hash used by Memory retrieval | Generated or model-written summaries are never copied into `docstring` or `source_docstring`, and `/ingest-code` never rewrites source files to add docstrings. Authored docstrings are preferred in retrieval text. A derived summary may appear only as `derived_summary.status="derived_unreviewed"` and only while its source/evidence hashes match the current symbol version; stale or malformed summaries fail closed to `null`. Identifier-heavy fields are emitted as `lexical_terms` such as `symbol:build_evidence_case`, `param:enable_llm`, `call:execute_llm_request`, and split identifier tokens. These are inputs to `/memory`'s code retrieval backend; `/ingest-code` does not create Qdrant collections or payload indexes directly. ### Typed Code Edges The local `edges.jsonl` bundle uses typed `CodeEdgeRecord` entries for file and symbol relationships: | Field | Purpose | |-------|---------| | `from_id`, `from_entity_type`, `to_id`, `to_entity_type` | Stable file/symbol endpoints; resolved edges must point at records present in the same bundle | | `edge_type` | One of `DEFINES`, `IMPORTS`, `CALLS`, `INHERITS`, `IMPLEMENTS` | | `resolution_status` | `resolved`, `candidate`, or `unresolved`; also mirrored as legacy `status` | | `resolution_method`, `confidence`, `provenance`, `synthesized_by` | How the edge was produced and how strong the static resolution is | | `source_path`, `source_start_line`, `source_end_line`, `source_start_column`, `source_end_column` | Exact source occurrence span for the edge | | `active_for_traversal` | `true` only for resolved edges; candidates and unresolved references are never traversal-active | | `raw_reference`, `candidate_ids`, `candidate_descriptors`, `unresolved_reason`, `attempted_resolution_stages` | Audit data for alias, relative import, ambiguous dispatch, and reflection/dynamic call cases | Current Python support resolves exact file imports, relative imports, explicit import aliases, local lexical calls, inherited method calls, and same-module/package call targets. Same-named functions that cannot be disambiguated are emitted as inactive candidates. Dynamic/reflection calls such as `getattr(...)` are emitted as inactive unresolved references unless a later resolver can prove a concrete target. Only resolved legacy import dependencies are sent to `/memory /add-edges`. The local typed bundle is the provenance receipt; candidates and unresolved edges are retained there for review but are not admitted as canonical active graph edges. ## Directory Filtering **Git repositories:** When scanning a git repo, `/ingest-code` uses `git ls-files` which automatically respects `.gitignore`. Files ignored by git are excluded from ingestion. **Hardcoded skip directories** (always skipped, even in non-git dirs): `.venv`, `venv`, `node_modules`, `__pycache__`, `.git`, `dist`, `build`, `.eggs`, `.mypy_cache`, `.pytest_cache`, `site-packages`, `.uv` **Always included** (regardless of .gitignore): `CONTEXT.md`, `README.md`, `CLAUDE.md`, `MEMORY.md`, `AGENTS.md`, plus `docs/` and `local/docs/`. ## Integration with /monitor-codebase The nightly pipeline calls `rescan` with scoped directories from `.monitor-codebase.json`: ```json { "include_dirs": ["src/extractor/pipeline/steps", "prototypes/tabbed/api"], "exclude_dirs": [".venv", "node_modules", "checkpoints"] } ``` The `exclude_dirs` list is additive — it supplements both `.gitignore` and the hardcoded skip directories. ## Output Format ```json { "files_scanned": 968, "knowledge_extracted": 1547, "knowledge_stored": 1520, "files_with_cwes": 23, "total_cwe_mappings": 45, "cwe_stored": 45, "cwe_summary": {"CWE-78": 5, "CWE-20": 12} } ``` ## Indexing Marker After a successful scan, `/ingest-code` writes a `.ingest-code.json` marker file to the scanned directory: ```json { "ingested_at": "2026-04-14T13:30:00", "path": "${HOME}/workspace/my-project", "stem": "my-project", "files_scanned": 968, "knowledge_stored": 1520, "cwe_stored": 45, "edges_stored": 234, "code_index": { "enabled": true, "backend": "memory", "collection": "code_symbols", "treesitter": true, "symbols_stored": 4182, "lexical_terms": true, "line_ranges": true, "content_hashes": true, "hybrid_retrieval_capable": true }, "scope": "code" } ``` **Why this exists:** Other skills (like `/code-runner`) can check for this marker to determine if a codebase has been semantically indexed. If `code_index.enabled` is true, they should prefer `/memory recall` over `code_symbols`/hybrid code retrieval before falling back to ripgrep pattern matching. The marker is also stored in `/memory` with tags `["ingest-code", "indexed-codebase", <stem>, <path>]` for discovery via recall. ## Incremental Re-Indexing Re-ingesting a repo emits a complete desired bundle while parsing only files whose source or relevant transforms changed. The bundle remains the handoff authority; the local cache is disposable acceleration evidence, not canonical backend state. Two local state files may exist: - `artifacts/ingest-code/incremental-components.json` stores file components from the last accepted complete bundle: source fingerprint, explicit transform
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub