- name
- ingest-code
- description
- Ingest codebases into /memory for knowledge extraction and CWE scanning. Phase 1 extracts functional knowledge (module docstrings, function signatures, class hierarchies, markdown docs) via Python AST. Phase 2 scans for CWE mappings via /taxonomy. Designed to run nightly via /monitor-codebase.
- allowed-tools
- ["Bash","Read"]
- triggers
- ["ingest code","scan codebase","ingest codebase","scan for cwes","codebase ingestion","code to memory"]
- metadata
- {"short-description":"Codebase knowledge + CWE extraction to /memory","author":"Horus","version":"0.2.0"}
- provides
- ["ingest-code"]
- composes
- ["memory","taxonomy","task-monitor","agentic-evals"]
- disciplines
- ["data-engineering","memory-knowledge"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.
# ingest-code
Codebase ingestion into `/memory`:
1. **Phase 1 — Functional Knowledge**: Python AST extracts module docstrings, function signatures, class hierarchies. Markdown parser extracts section-level knowledge from CONTEXT.md, README.md, etc. Generic parser handles TS/JS exports.
2. **Phase 2 — CWE Scanning**: `/taxonomy` extracts security-relevant patterns (bridge tags + CWE mappings) per file.
3. **Phase 3 — Relationship Edges**: Python import analysis stores resolved code dependency edges in `/memory`; the local code-graph bundle also records typed import/call/inheritance occurrences, including inactive unresolved and ambiguous candidates for audit.
4. **Phase 4 — Structured Code Index**: Optional Tree-sitter extraction emits a complete code-graph bundle and applies it through Memory/GMO's governed `/code/projection/apply` lifecycle endpoint.
5. **Static Debugger Affordances**: Code-graph bundles emit `debugger.invocation_candidate.v1` rows in `debug_invocations.jsonl`. These are candidate routes only; `$debugger` must prove a candidate at runtime before Memory can promote it.
Functional lessons and CWE summaries remain lesson-style memory records for compatibility. Structured files, symbols, and edges are canonicalized by Memory/GMO's complete projection lifecycle, not by independent per-symbol batches. `/memory` owns ArangoDB, Qdrant, embeddings, sparse/hybrid retrieval, and payload/index behavior. `/ingest-code` must not talk to ArangoDB or Qdrant directly.
## Quick Start
```bash
cd .pi/skills/ingest-code
# Full knowledge + CWE scan
./run.sh scan /path/to/codebase
# CWE scan only (legacy mode)
./run.sh scan /path/to/codebase --cwe-only
# Preview without storing to /memory
./run.sh scan /path/to/codebase --dry-run
# Include Tree-sitter structured code symbols for memory's hybrid code index
./run.sh scan /path/to/codebase --treesitter
# Nightly rescan (only files modified in last day)
./run.sh rescan --since 1d -c /path/to/codebase --treesitter
# Read-only target freshness check before using Memory code snippets for repair
./run.sh ensure-current \
--repo /path/to/codebase \
--branch main \
--commit "$COMMIT_SHA" \
--path src/example.py \
--json
```
## Commands
### `scan` — Full Codebase Scan
```bash
./run.sh scan <path> [OPTIONS]
Options:
--glob, -g File patterns (default: *.py *.ts *.js *.rs *.go *.java *.c *.cpp)
--cwe-only Skip Phase 1, only run CWE scan
--validate Run LLM validation on CWE matches
--treesitter Run Tree-sitter scan for structured code symbols
--code-index Apply the complete Tree-sitter code graph bundle to Memory/GMO (default)
--no-code-index Disable structured code projection application
--compat-symbol-upsert
Use legacy per-symbol Memory upserts with a visible warning
--dry-run Preview without writing to /memory
--scope Memory scope (default: "code")
--batch-size Files per CWE scan batch (default: 50)
```
### `rescan` — Incremental Rescan (Scheduler Job)
```bash
./run.sh rescan [OPTIONS]
Options:
--since Only files modified since (ISO date or "1d", "7d")
-c, --codebase Codebase path(s) to rescan (repeatable)
--validate Run LLM validation
--treesitter Run Tree-sitter scan for structured code symbols
--code-index Apply the complete Tree-sitter code graph bundle to Memory/GMO (default)
--no-code-index Disable structured code projection application
--scope Memory scope
```
### `ensure-current` — Target-Scoped Projection Freshness Preflight
```bash
./run.sh ensure-current [OPTIONS]
Options:
--repo Repository worktree to check (required)
--branch Expected branch/ref name (default: current branch)
--commit Expected commit SHA (default: current HEAD)
--path Repository-relative target path; repeatable
--scope Memory/GMO projection scope (default: "code")
--json Emit `ingest-code.code_projection_freshness.v1`
--refresh Explicitly refresh through `scan --treesitter --code-index`
--canonical-branch Branch allowed to activate canonical projection (default: "main")
--max-target-files Bound directory expansion (default: 200)
```
`ensure-current` is the pre-repair gate for stateless workers. It resolves the
repository root, branch, commit, and target paths; rejects absolute paths,
`..`, and repository escapes; reads active code-search/code-node/code-coverage
state through the supported Memory/GMO code-navigation boundary; then compares
current source hashes with indexed source hashes for the requested targets.
The result status is one of:
| Status | Meaning |
|--------|---------|
| `CURRENT` | Active Memory/GMO source hashes match current target files and coverage allows modification guidance. |
| `SOURCE_CURRENT_INDEX_INCOMPLETE` | Source bytes match, but coverage is incomplete, so exhaustive callers/callees/impact absence claims are blocked. |
| `STALE` | Current source differs from the active projection; stored snippets are not modification authority. |
| `UNINDEXED` | No applicable active projection record matched the target. |
| `BLOCKED` | Identity, containment, service, receipt, or validation failed closed. |
Default `ensure-current` is read-only. It must not parse files, create
embeddings, apply projections, or fall back to legacy per-symbol writes. With
`--refresh`, it may run the existing complete-bundle scan only when the checkout
is clean, on the configured canonical branch, and bound to the requested
commit. Feature/repair worktrees are refused so unreviewed code cannot replace
the canonical main projection.
## What Gets Extracted
### Phase 1: Functional Knowledge (Python AST)
| Source | What | Example /memory Problem |
|--------|------|------------------------|
| Module docstring | Module purpose | "What does run_pipeline.py do?" |
| Class definition | Class + methods + bases | "What is the ContentRepository class in content_query.py?" |
| Function signature | Args, return type, docstring | "What does extract_tables() do in s05_table_extractor.py?" |
| Markdown sections | Architecture decisions, bug fixes | "What does 'Bugs Fixed' say in MEMORY.md?" |
| TS/JS exports | Exported symbols | "What is AnswerCanvas in AnswerCanvas.tsx?" |
### Phase 2: CWE Scanning (via /taxonomy)
| Category | Example CWEs | Triggers |
|----------|--------------|----------|
| MemorySafety | CWE-120, CWE-787, CWE-416 | buffer, overflow, memory, pointer |
| InputValidation | CWE-20, CWE-89, CWE-78 | input, validation, inject, command |
| Authentication | CWE-287, CWE-798, CWE-522 | auth, credential, password, session |
| Cryptography | CWE-311, CWE-327, CWE-330 | encrypt, crypto, key, random |
### Phase 4: Structured Code Index (Tree-sitter → /memory)
When `--treesitter --code-index` is enabled, `/ingest-code` writes a deterministic local code-graph bundle under `artifacts/ingest-code/code-graph/`, computes the submitted bundle digest and checksums digest, and submits the complete bundle to Memory/GMO through `/code/projection/apply`.
The Memory/GMO receipt is the only canonical projection success signal. It must bind the submitted bundle digest, checksums digest, activated generation, and expected file/symbol/edge counts. If Memory/GMO is unavailable, rejects the bundle, or returns a receipt whose digest does not match the submitted bundle, the scan fails closed and does not fall back to per-symbol writes.
Each record includes:
| Field | Purpose |
|-------|---------|
| `repo`, `branch`, `commit`, `path` | Scope and staleness control |
| `language`, `symbol_kind`, `symbol_name`, `qualified_name` | Symbol filtering and exact lookup |
| `start_line`, `end_line`, `code`, `content_hash` | Cited source retrieval and deterministic updates |
| `imports`, `parameters`, `local_variables`, `called_symbols`, `string_literals` | Lexical terms for memory's sparse/hybrid retrieval |
| `problem`, `solution`, `text`, `tags` | Compatibility with existing memory recall surfaces |
Documentation metadata is provenance-safe:
| Field | Purpose |
|-------|---------|
| `source_docstring` / `docstring` | Exact authored source docstring text, preserved for compatibility |
| `source_docstring_status` | `present`, `missing`, `generated_file`, or `not_applicable` for v1 extraction |
| `documentation_need` | Deterministic triage: `required`, `recommended`, `optional`, or `exempt` |
| `documentation_need_reasons` | Source-derived reasons such as `public_api`, `external_io`, `security`, `mutation`, or `trivial_helper` |
| `summary_evidence` | Canonical source-fact packet and hash for optional generated summaries |
| `derived_summary` | Current unreviewed generated summary only when bound to the current `symbol_version_id`, source hash, and evidence hash |
| `retrieval_text`, `retrieval_text_sha256`, `purpose_source` | Canonical semantic text and hash used by Memory retrieval |
Generated or model-written summaries are never copied into `docstring` or
`source_docstring`, and `/ingest-code` never rewrites source files to add
docstrings. Authored docstrings are preferred in retrieval text. A derived
summary may appear only as `derived_summary.status="derived_unreviewed"` and
only while its source/evidence hashes match the current symbol version; stale
or malformed summaries fail closed to `null`.
Identifier-heavy fields are emitted as `lexical_terms` such as `symbol:build_evidence_case`, `param:enable_llm`, `call:execute_llm_request`, and split identifier tokens. These are inputs to `/memory`'s code retrieval backend; `/ingest-code` does not create Qdrant collections or payload indexes directly.
### Typed Code Edges
The local `edges.jsonl` bundle uses typed `CodeEdgeRecord` entries for file and symbol relationships:
| Field | Purpose |
|-------|---------|
| `from_id`, `from_entity_type`, `to_id`, `to_entity_type` | Stable file/symbol endpoints; resolved edges must point at records present in the same bundle |
| `edge_type` | One of `DEFINES`, `IMPORTS`, `CALLS`, `INHERITS`, `IMPLEMENTS` |
| `resolution_status` | `resolved`, `candidate`, or `unresolved`; also mirrored as legacy `status` |
| `resolution_method`, `confidence`, `provenance`, `synthesized_by` | How the edge was produced and how strong the static resolution is |
| `source_path`, `source_start_line`, `source_end_line`, `source_start_column`, `source_end_column` | Exact source occurrence span for the edge |
| `active_for_traversal` | `true` only for resolved edges; candidates and unresolved references are never traversal-active |
| `raw_reference`, `candidate_ids`, `candidate_descriptors`, `unresolved_reason`, `attempted_resolution_stages` | Audit data for alias, relative import, ambiguous dispatch, and reflection/dynamic call cases |
Current Python support resolves exact file imports, relative imports, explicit import aliases, local lexical calls, inherited method calls, and same-module/package call targets. Same-named functions that cannot be disambiguated are emitted as inactive candidates. Dynamic/reflection calls such as `getattr(...)` are emitted as inactive unresolved references unless a later resolver can prove a concrete target.
Only resolved legacy import dependencies are sent to `/memory /add-edges`. The local typed bundle is the provenance receipt; candidates and unresolved edges are retained there for review but are not admitted as canonical active graph edges.
## Directory Filtering
**Git repositories:** When scanning a git repo, `/ingest-code` uses `git ls-files` which automatically respects `.gitignore`. Files ignored by git are excluded from ingestion.
**Hardcoded skip directories** (always skipped, even in non-git dirs):
`.venv`, `venv`, `node_modules`, `__pycache__`, `.git`, `dist`, `build`, `.eggs`, `.mypy_cache`, `.pytest_cache`, `site-packages`, `.uv`
**Always included** (regardless of .gitignore): `CONTEXT.md`, `README.md`, `CLAUDE.md`, `MEMORY.md`, `AGENTS.md`, plus `docs/` and `local/docs/`.
## Integration with /monitor-codebase
The nightly pipeline calls `rescan` with scoped directories from `.monitor-codebase.json`:
```json
{
"include_dirs": ["src/extractor/pipeline/steps", "prototypes/tabbed/api"],
"exclude_dirs": [".venv", "node_modules", "checkpoints"]
}
```
The `exclude_dirs` list is additive — it supplements both `.gitignore` and the hardcoded skip directories.
## Output Format
```json
{
"files_scanned": 968,
"knowledge_extracted": 1547,
"knowledge_stored": 1520,
"files_with_cwes": 23,
"total_cwe_mappings": 45,
"cwe_stored": 45,
"cwe_summary": {"CWE-78": 5, "CWE-20": 12}
}
```
## Indexing Marker
After a successful scan, `/ingest-code` writes a `.ingest-code.json` marker file to the scanned directory:
```json
{
"ingested_at": "2026-04-14T13:30:00",
"path": "${HOME}/workspace/my-project",
"stem": "my-project",
"files_scanned": 968,
"knowledge_stored": 1520,
"cwe_stored": 45,
"edges_stored": 234,
"code_index": {
"enabled": true,
"backend": "memory",
"collection": "code_symbols",
"treesitter": true,
"symbols_stored": 4182,
"lexical_terms": true,
"line_ranges": true,
"content_hashes": true,
"hybrid_retrieval_capable": true
},
"scope": "code"
}
```
**Why this exists:** Other skills (like `/code-runner`) can check for this marker to determine if a codebase has been semantically indexed. If `code_index.enabled` is true, they should prefer `/memory recall` over `code_symbols`/hybrid code retrieval before falling back to ripgrep pattern matching.
The marker is also stored in `/memory` with tags `["ingest-code", "indexed-codebase", <stem>, <path>]` for discovery via recall.
## Incremental Re-Indexing
Re-ingesting a repo emits a complete desired bundle while parsing only files
whose source or relevant transforms changed. The bundle remains the handoff
authority; the local cache is disposable acceleration evidence, not canonical
backend state.
Two local state files may exist:
- `artifacts/ingest-code/incremental-components.json` stores file components
from the last accepted complete bundle: source fingerprint, explicit transform
View on GitHub