- name
- aim-sot
- description
- Tracks the source-of-truth for each part of the user's own project — registry schema, templates, and engine (consult / detect-propose / verify). Detects file and directory drift (content-digest and directory tree-digest), runs a machine-local shadow-git history for change detection, and correlates code changes to stale docs (doc-drift via DOCOWNERS). Use when checking SOT drift, choosing a drift_strategy, setting up the shadow git, or reviewing doc-staleness findings. Do NOT use for AI Memory search or save operations, or for checking AI Memory system health (use aim-search, aim-save, or aim-status).
- allowed-tools
- Bash, Read
# aim-sot — Source-of-Truth Subsystem
Manages a user-committed `.sot/registry.yaml` that declares where the canonical truth lives for each boundary of the user's project. Every entry is a pointer + provenance record — no copied source content ever.
## Overview
The registry lives in **the user's own repository** at `.sot/registry.yaml`, committed alongside code and diff-reviewable by the team. The skill ships the schema, templates, and engine. No project-specific data is baked into the skill.
## Modes
Three engine modes:
- **consult** — read-only query engine over the user's committed `.sot/registry.yaml`. Subcommands: `list` (all entries), `get <id>` (full entry), `where <id>` (sot_location), `who <id>` (owner), `drift <id>` (drift_check). Global flags: `--registry PATH` (override path), `--json` (machine-readable output). Invoked via `run-with-env.sh` (Pattern B, BP-013). Script: `_ai-memory/skills/aim-sot/scripts/aim_sot_consult.py`.
- **detect-propose** — hybrid auto-discover → propose: scans for candidate components, computes actual state, compares to the registry, and emits a proposed patch on drift or new candidates. Never writes the registry directly.
- **verify** — 16-check gate (Schema · Referential · Completeness · Content). Mandatory before any apply; human approval (HITL) required. CI/pre-commit hook is opt-in only, never auto-installed.
## Lifecycle
### Create (bootstrap a new project)
1. **Discover** — run `detect-propose run` (see **Detect-Propose — Invocation**); emits candidate proposals, never writes the registry.
2. **Author** — apply the [**Authoring**](#authoring) loop to each kept candidate; assemble into `.sot/registry.yaml` starting from `templates/registry.yaml.template`.
3. **Verify** — run `verify` (see **Verify — Invocation**); **0 failures required**. On a fresh registry the verdict is **CONDITIONAL, not PASS** — cold-start K1 `skipped_no_baseline` and advisory (C1) warnings are expected and do not block apply (see **Verdicts**); literal `PASS` needs an established baseline.
4. **Apply** — human applies the registry after a clean verdict (HITL).
### Update (ongoing — drift and new candidates)
1. **Detect** — run `detect-propose run`; emits drift proposals and new-candidate proposals; never writes the registry.
2. **Author** — apply the [**Authoring**](#authoring) loop to any changed or new entries.
3. **Gate** — run `verify` via the **pre-apply proposal gate** (see **Usage — pre-apply proposal gate** under **Verify — Invocation**); mandatory before applying.
4. **Apply** — human applies; drift is never auto-applied.
### Cross-cutting invariants
- **Propose-only** — `detect-propose` never writes `.sot/registry.yaml`.
- **Verify gates apply** — the 16-check `verify` gate is mandatory before any apply.
- **Human applies** — every registry change requires explicit human action (HITL).
- **Never auto-rewrite** — no mode auto-applies drift or candidate proposals.
## Consult — Invocation
```bash
bash "${AI_MEMORY_INSTALL_DIR:-$HOME/.ai-memory}/scripts/memory/run-with-env.sh" \
"${AI_MEMORY_INSTALL_DIR:-$HOME/.ai-memory}/_ai-memory/skills/aim-sot/scripts/aim_sot_consult.py" \
<subcommand> [--json] [--registry PATH]
```
## Detect-Propose — Invocation
```bash
bash "${AI_MEMORY_INSTALL_DIR:-$HOME/.ai-memory}/scripts/memory/run-with-env.sh" \
"${AI_MEMORY_INSTALL_DIR:-$HOME/.ai-memory}/_ai-memory/skills/aim-sot/scripts/aim_sot_detect_propose.py" \
run [--json] [--registry PATH] [--limit N] [--all] [--write-proposal] [--force]
```
```bash
# Rebuild the 5b derived memory cache explicitly:
bash "${AI_MEMORY_INSTALL_DIR:-$HOME/.ai-memory}/scripts/memory/run-with-env.sh" \
"${AI_MEMORY_INSTALL_DIR:-$HOME/.ai-memory}/_ai-memory/skills/aim-sot/scripts/aim_sot_detect_propose.py" \
reindex [--registry PATH]
```
**Hard invariant**: `detect-propose` NEVER writes `.sot/registry.yaml` — proposals only.
The human applies edits manually; the registry is always committed and diff-reviewable.
### First run — bootstrap from zero
When no `.sot/registry.yaml` exists yet, `detect-propose run` still runs the
discovery scan and emits **candidate proposals** for the project (it does not bail).
Bootstrapping a new project:
1. **Discover** — run `detect-propose run` (optionally `--all` to lift the cap). The
scan roots at the project root, or the current working directory when no registry
is present, and prints discovered candidates. Add **`--write-proposal`** to also
scaffold a ready-to-edit staging draft (see below).
2. **Author** — for each candidate you keep, apply the [**Authoring**](#authoring) loop
(categorize by type → propose entries → write each description to the D1–D4 rubric →
fix any FAIL before emit); assemble the results into a new `.sot/registry.yaml`
starting from `templates/registry.yaml.template` (or from the `--write-proposal`
draft).
3. **Verify** — run `aim-sot verify` to gate the registry through the 16-check taxonomy.
4. **Approve** — apply after a clean verdict + human approval (HITL).
Discovery is still **propose-only** with no registry: candidates go to stdout and the
committed registry is never created or written.
#### `--write-proposal` — scaffold a staging draft (TD-744)
`detect-propose run --write-proposal` writes a **non-committed** staging file
`.sot/registry.proposed.yaml` from the discovered candidates, so you edit a
schema-shaped draft instead of transcribing a stdout list by hand. It fires **only
on a registry-less project** (the entry point that already runs cold-start
discovery); when a committed `.sot/registry.yaml` exists the registry-present drift
path takes over and `--write-proposal` does nothing.
- **What it fills.** Structural fields are written as real values (`boundary_type`,
`sot_location`, `status: proposed`, `added_by: "aim-sot bootstrap"`). Every
human-owned semantic field is an explicit `TODO(human): …` placeholder
(`kind`, `owner`, `description`, `last_verified`, `provenance_note`) — **never
auto-filled (BP-029)**. Each entry's `confidence` + mandatory `inferred_from` ride
as advisory comments above it (they are not registry-schema fields).
- **Promotion = approval.** The engine writes a *draft*, never the SoT. You fill the
`TODO(human)` fields, prune candidates, then **rename** the draft to
`.sot/registry.yaml`, commit it, and run `aim-sot verify`. That promotion + commit
is the human approval (BP-030). Recommend git-ignoring `.sot/registry.proposed.yaml`.
- **Idempotency.** A second run **skips** an existing draft (to protect in-progress
edits) and says so; pass **`--force`** to overwrite. The committed registry is never
written under any path (BP-030 invariant).
- **Owner candidates (advisory).** When the project is a git repo, the draft carries
an `owner_candidates` advisory comment block — the top-3 committers per
`sot_location` from `git log`, as a hint for authoring `owner` by hand. It is
**never** written into the `owner` field, and **degrades silently** (block omitted)
on a non-git project.
#### Confidence tiers (TD-744 Q5)
Candidates carry an ordinal `confidence` triage label paired with `inferred_from`:
| Tier | Signal | Example `inferred_from` |
|------|--------|-------------------------|
| `high` | the project *declares* the unit | a manifest filename, `adr_directory` |
| `medium` | structural heuristic (a top-level folder) | `top_level_directory` |
| `low` | weak/ambiguous (nested source dir, no manifest) | `nested_source_directory` |
The label is **review-triage priority only** — it is never an auto-approve gate and
never licenses auto-filling semantics; even a `high` candidate needs human authoring.
### Flags (`run`)
| Flag | Default | Description |
|------|---------|-------------|
| `--limit N` | 20 | Cap new-candidate proposals per run |
| `--all` | off | Disable cap — surface all new candidates |
| `--json` | off | Machine-readable JSON output |
| `--registry PATH` | (git root) | Override registry path |
| `--write-proposal` | off | Registry-less project only: scaffold a non-committed `.sot/registry.proposed.yaml` staging draft (structural fields filled, semantic fields `TODO(human)`). Never writes the committed registry. See [First run — bootstrap from zero](#first-run--bootstrap-from-zero). |
| `--force` | off | With `--write-proposal`, overwrite an existing draft (default: skip to protect in-progress edits). |
| `--shadow` | off | Run the `[CL]` detect pass — BP-039 digest → BP-040 shadow-git commit → BP-042 doc-drift → structured `findings`. The Stop hooks pass it; see [Drift strategies, shadow git & doc-drift](#drift-strategies-shadow-git--doc-drift). |
| `--drift-only` | off | Hot path only: check the registered entries for drift and skip the discovery walk entirely (BP-047). Mutually exclusive with `--discover`. **Does not skip the 5b reindex gate** — a registry-content change still triggers the derived-cache rebuild; only `--no-reindex` suppresses Qdrant writes. |
| `--discover` | off | Force a full discovery walk now, bypassing the periodic cadence gate (TTL / session-count) and resetting it (BP-047). Mutually exclusive with `--drift-only`. |
| `--no-reindex` | off | Skip the 5b derived-cache reindex so drift / discovery can be exercised with **zero** Qdrant writes (offline / sandbox). |
### Output — drift proposals
Each drift proposal includes `drift_type` (one of `location / staleness_temporal /
staleness_hash / declaration_reality`), `root_cause`, `impact`, and
`recommended_action`. A `declaration_reality` entry with `"k1_trigger": true`
requires mandatory human re-confirmation before the description is considered valid
(spec §5 K1).
**Dual-firing on content change (by design):** when a file's hash changes, the engine
emits *both* `staleness_hash` (re-verify the artifact) *and* `declaration_reality`
(K1 trigger: description/tags must be re-confirmed by a human). These are
complementary signals — consumers may act on `k1_trigger` independently of the
staleness signal. Do not treat the two as duplicates; they require different actions.
**Drift digest is dispatched by strategy:** a **file** `sot_location` uses
`content-digest` (sha256 of the file — unchanged behavior); a **directory**
`sot_location` uses `tree-digest` (BP-039 sorted-per-file-SHA-256 over the tree),
so directory boundaries now get content-hash drift too. Override per entry with the
schema-validated `drift_strategy` enum (never a shell command). A strategy switch or
digest-version bump **re-baselines** the entry rather than firing drift. See
[Drift strategies, shadow git & doc-drift](#drift-strategies-shadow-git--doc-drift).
### Output — new candidate proposals
Structural fields (`boundary_type`, `sot_location`, `confidence`, `inferred_from`)
are auto-inferred. **Semantic fields (`owner`, `description`, `provenance_note`)
are never auto-filled — human-authored always (BP-029).**
### 5a and 5b caches
- **5a** (`~/.ai-memory/drift-state/sot_drift_{project_id}.json`) — per-install drift
state; never committed.
- **5b** (Qdrant `conventions` collection, `type=sot_entry`) — derived memory
cache; deterministically rebuildable from the committed registry via `reindex`.
#### Connect to the SOT / memory Qdrant from a scratch script
To inspect the 5b collection (e.g. verify `--no-reindex` wrote nothing), use the
project's client factory rather than constructing `QdrantClient` by hand — it
resolves host/port, the API key, HTTPS, and the gRPC-vs-HTTP transport from config,
avoiding the "illegal request line" / TLS `WRONG_VERSION_NUMBER` errors an ad-hoc
client hits:
```python
import os, sys
sys.path.insert(0, os.path.expanduser("~/.ai-memory/src")) # the installed runtime
from qdrant_client import models
from memory.qdrant_client import get_qdrant_client
client = get_qdrant_client(read_only=True) # defaults to prefer_grpc=True
count = client.count(
"conventions",
# gRPC mode requires a typed Filter, not a raw dict.
count_filter=models.Filter(
must=[
models.FieldCondition(
key="type", match=models.MatchValue(value="sot_entry")
)
]
),
).count
print("sot_entry records:", count)
```
## Drift strategies, shadow git & doc-drift
The `--shadow` flag on `detect-propose run` runs the `[CL]` detect pass: one BP-039
tree-digest of the project → on change, a machine-local **shadow-git** commit
(BP-040) → a `git diff` → **doc-drift** correlation (BP-042) → a structured
**findings** pipe. All of it is engine-side, shared by every CLI; the per-CLI Stop
hooks are thin callers that pass `--shadow`.
Key guarantees:
- **Non-invasive** — the shadow git is a bare repo under `~/.ai-memory/sot-git/<project_id>/` driven by an explicit two-pointer (`GIT_DIR`/`GIT_WORK_TREE`); it writes **nothing** into the user's tree (no `.git`, no `.gitignore`). Teardown is `rm -rf`.
- **git-required, project-need-not-be-git** — git is a required tool, but the portable directory tree-digest never needs the project itself to be a git repo.
GitHub에서 보기