- name
- memory
- description
- MEMORY FIRST - Query memory BEFORE scanning any codebase. Use when encountering ANY problem, error, or task. Call "recall" FIRST, then scan codebase only if nothing found. Triggers: "check memory", "recall", "have we seen", "remember how".
- allowed-tools
- Bash, Read
- triggers
- ["assess memory usage","check memory API usage","check memory","recall","clarify","have we seen this","remember how we solved","what did we learn","recall previous","save this lesson","learn from this","check memory for","have we seen this before","query memory first","ask clarifying questions","doesn't understand"]
- metadata
- {"short-description":"MEMORY FIRST - Query before scanning codebase"}
- provides
- ["memory-recall","memory-learn","intent-classification","edge-verification","usage-assessment"]
- composes
- ["extractor","edge-verifier","taxonomy","embedding","task-monitor","agentic-evals"]
- complies
- ["best-practices-skills","best-practices-python","best-practices-scillm","best-practices-arangodb"]
- taxonomy
- ["knowledge","persistence","resilience","precision"]
- docs
- {"arangodb":"/best-practices-arangodb"}
- disciplines
- ["memory-knowledge"]
> **STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.**
> Do not skim. Do not skip to the code examples. This document contains
> unique constraints, deterministic `_key` rules, secondary unique indexes,
> Qdrant semantic-sync metadata rules (no Arango vector arrays), and schema ownership rules that
> WILL cause silent data corruption or 400/409 errors if you ignore them.
> Every section exists because an agent broke something by not reading it.
# Memory Skill - MEMORY FIRST Pattern
Pi is the only CLI agent that can reliably enforce Memory First (other CLIs treat pre-hooks as optional), so this skill is the **front-door contract** for Pi and humans alike.
**Non-negotiable rule**: Query memory BEFORE scanning any codebase.
## Commands Snapshot
| Command | Use Case |
| ------------------------------------------------- | ---------------------------------------------------------------------- |
| `./run.sh recall --q "..." --brief` | **DEFAULT.** Slim output + proven skill chain. Use this. |
| `./run.sh recall --q "..."` | Full output with taxonomy, raw scores, _key (when you need metadata) |
| `./run.sh intent --q "..."` / `httpx POST /intent {q, scope?, fast?, app?}` | First-class intent product: classify route, QuerySpec, `recall_profile`, slots, and query plan |
| `httpx POST /speaker/resolve {candidates, threshold?, persona_id?}` | Voice identity front door: resolve listener evidence into known/unknown/ambiguous speaker context before `/intent` or personal recall |
| `httpx POST /answer {q, scope?, k?, collections?}` | First-class grounded answer product: answer only from deterministic general answers or recall evidence |
| `./run.sh clarify --q "..."` / `httpx POST /clarify {q, scope?, context?, k?}` | First-class ambiguity product: ask targeted follow-up questions when the request is underspecified |
| `./run.sh deflect --q "..."` / `httpx POST /deflect {q, persona_id?, intent_action?}` | First-class deflection product: route off-topic, unsafe, or no-match turns away from memory/evidence work |
| `httpx POST /execution-runs {...duration fields...}` | Record actual route/tool/subagent execution duration for ETA learning and recall |
| `httpx POST /execution-stats {route_key? ...}` | Return p50/p75/p90/p95 duration estimates from stored `execution_runs` |
| `httpx POST /store {document, collection}` | **THE write endpoint.** Write to ANY collection. Auto-upserts by `_key`. |
| `httpx POST /store {document}` (no collection) | Writes to `lessons` with Qdrant semantic sync + dedup (same as old `/learn`) |
| `httpx POST /upsert {collection, documents}` | Batch write (multiple docs). Same rules as `/store`. |
| `memory-agent activity ingest-git REPO ...` | Ingest git commits as durable `project_activity` records for day/week code-activity recall |
| `uv run python -m graph_memory.maintenance.sanity_recall persona-graph-materialize` | Materialize searchable persona relationship docs into true Arango graph edge collections |
| `uv run python -m graph_memory.maintenance.sanity_recall persona-graph --base-url http://127.0.0.1:8601` | E2E canary: natural-language persona recall returns source/evidence-grounded memories, QRA-style question/answer/evidence/source records, explicit 2+ hop graph traversal, and directly filterable `tom_edges` |
| `./run.sh learn --problem "..." --solution "..."` | **Deprecated.** CLI shorthand that calls `/store` with `collection=lessons` |
| `./run.sh chain-recall "query"` | Search proven skill chains directly |
| `./run.sh chain-learn --skills "a,b,c" --task "..."` | Store a proven skill chain |
| `./run.sh chain-stats` | Skill chain collection statistics |
| `./run.sh preset compile --ids '{"set":"..."}'` | Compile deterministic technical specs from ArangoDB |
| `./run.sh preset sanity` | Audit preset library for broken links / cycles (Strict Mode) |
| `./run.sh info` | Print active configuration (embedder, episodic sources, edge verifier) |
| `./run.sh serve --host --port` | Keep the FastAPI server warm for low-latency recall |
| `./run.sh status` | Quick health check / Arango connectivity |
## `--brief` Mode: Context-Safe Recall with Skill Chains
**Use `--brief` by default.** It returns ~3.5x smaller output (problem, solution,
score, tags) PLUS the best matching proven skill chain from the `skill_chains`
collection. This is the "have I solved this before, and what skills worked?" pattern.
```bash
./run.sh recall --q "checkpoint resume fails after clear" --brief
```
```json
{
"found": true,
"items": [
{
"problem": "checkpoints collection not searchable via /recall",
"solution": "Added to builtin_sources() in _declarations.py...",
"score": 0.99,
"tags": ["checkpoint", "grade:clean"]
}
],
"skill_chain": {
"skills": ["memory", "assess", "checkpoint"],
"task_type": "general",
"success_rate": 1.0,
"observations": 3,
"elegance": "efficient",
"score": 0.78
}
}
```
**If `skill_chain` is present: follow it.** These chains are extracted from real
commits across 11 repos (17K+ commits) and proven by successful outcomes. The
agent doesn't guess which skills to compose — it follows the proven path.
### How Skill Chains Are Built
```
/checkpoint --skills A B C --grade clean
↓
1. Git commit with Skills: trailer (machine-readable)
2. learn_chain() → skill_chains collection (embeddings, energy scoring)
3. Nightly: mine-transcripts → chain-backfill → new chains from history
↓
Next agent: recall --brief → skill_chain: [A, B, C]
```
**Sources** (2,300+ chains, ranked by quality):
- `production` — from /checkpoint --skills (highest confidence)
- `commit-trailer` — from git commit Skills: trailers
- `commit-transcript` — transcript scan within ±15min of commit timestamp
- `transcript` / `warm_pond` — nightly regex-mined (lower confidence)
### Chain Prioritization
`--brief` prefers production chains over transcript-mined chains, and filters
out noisy chains with >8 skills. If no production chain matches, falls back
to transcript-mined chains that match semantically.
## Daemon HTTP Endpoints
Memory service runs as a Docker container on `http://127.0.0.1:8601`.
### First-Class Routing Products: Intent, Answer, Clarify, Deflect
`/intent`, `/answer`, `/clarify`, and `/deflect` are structured JSON products,
not prose helpers. They are a shared fail-closed
routing contract for memory, `/create-evidence-case`, SciLLM-backed final
responses, and delegated subagents. Do not let each skill invent its own
threshold for answerability, ambiguity, or off-topic rejection.
Use the routing set this way:
```text
/speaker/resolve for voice turns
-> /intent with speaker_resolution when speaker is known
-> /clarify when speaker is unknown or ambiguous
/intent
-> /answer when the request is grounded enough to answer
-> /clarify when entities, scope, evidence, or relationships are ambiguous
-> /deflect when the turn is off-topic, unsafe, no-match, or outside memory scope
```
These endpoints are routing or final-response products, not raw retrieval:
| Product | When to use | Schema | Surface |
| --- | --- | --- | --- |
| `/speaker/resolve` | Resolve listener/diarization/speaker-verification evidence into known, unknown, or ambiguous speaker context before using personal memory. It does not compute embeddings or inspect raw audio. | `memory.speaker_resolution.v1` | HTTP |
| `/intent` | Classify the user turn into a route, QuerySpec, recall profile, extracted entities, tag families, confidence/ranked candidates, slots, required artifacts, query plan, and turn-scoped `delivery_context` before retrieval or final-response work. | intent response fields + `memory.delivery_context.v1` | `./run.sh intent` and HTTP |
| `/answer` | Return clean grounded final text from deterministic general answers or memory recall evidence, plus engine-neutral `delivery_plan`. It must not invent unsupported facts or inject renderer tags. | `memory.answer.v1` + `memory.delivery_plan.v1` | HTTP only for now |
| `/clarify` | Ask targeted clean follow-up text when the query is too vague, has weak recall, unsupported entities, taxonomy gaps, or ambiguous scope, plus engine-neutral `delivery_plan`. | `memory.clarify.v1` + `memory.delivery_plan.v1` | `./run.sh clarify` and HTTP |
| `/deflect` | Redirect off-topic, unsafe, no-match, or content-safety turns before they enter recall, evidence-case, QRA, or subagent work, returning clean text plus engine-neutral `delivery_plan`. | `memory.deflect.v1` + `memory.delivery_plan.v1` | `./run.sh deflect` and HTTP |
#### Mandatory Hardening Trace And Human Checkpoints
Memory hardening runs must emit a durable per-question pipeline trace. This is
not optional debug garnish. If a hardening case has no trace, the case is
`NOT_ESTABLISHED` even when a final route was returned.
Each trace must include:
- the exact `question`;
- ordered `pipeline_steps[]` for `/intent`, entity extraction, policy gates,
recall/profile selection, `/create-evidence-case` when used, and terminal
`/answer`, `/clarify`, `/deflect`, or `/draft`;
- per-step `duration_ms`, `status`, and `result_summary`;
- grounded entity evidence such as `entities`, `valid_entities`,
`invalid_terms`, `frameworks`, and the grounding source;
- retrieval evidence summaries, including QRA source/admission state and any
display-only question similarity when a QRA is chosen;
- the terminal `final_action`;
- a `human_checkpoint` object whenever the terminal action is `clarify` or
`draft`;
- the receipt or artifact path that preserves the trace.
`human_checkpoint` is mandatory for `CLARIFY` and `DRAFT` routes. For
`CLARIFY`, it must expose the exact clarifying question, the grounded context
that caused clarification, and the specific human response needed to continue.
For `DRAFT`, it must expose editable `question`, `reasoning`, answer text,
parallel question variants when present, evidence/source basis, signoff state,
and the allowed human actions: accept, reject, amend, or request another draft.
Unsigned drafts are not answer authority.
The hardening loop is:
```text
question -> trace each Memory pipeline step -> terminal action
-> ANSWER/DEFLECT: record trace and continue the same failure family
-> CLARIFY: ask the human the exact clarification question, then resume with the response
-> DRAFT: ask the human to accept/reject/amend/request edit, then record signoff/rejection
```
Project agents must not keep iterating silently after a `CLARIFY` or `DRAFT`
trace. Those routes are collaboration checkpoints. If the same failure family
survives two focused repair attempts, stop solo patching. Use `$debugger` when
the next edit depends on live runtime state, and show the human the concrete
breakpoint, source line, and relevant paused values. If the problem is semantic
or policy judgment rather than hidden runtime state, ask the human the specific
question shown by the trace.
`/create-evidence-case` depends on this boundary: `ANSWER` means evidence is
coherent enough to synthesize; `CLARIFY` means the case should not force a
verdict yet; `DEFLECT` means the request should not enter the evidence pipeline.
SciLLM may write the human-facing `final_response`, but deterministic memory
logic owns the route state and source packet. Every SciLLM final-response call
from memory must include `X-Caller-Skill: memory` and source metadata.
#### Delivery Context, Emotion, And Realtime Voice Boundary
Memory owns grounded text and engine-neutral delivery metadata. Realtime voice
renderers such as SPARTA/Chatterbox own render manifests, injectable tags, and
playback parameters.
Do not put Chatterbox tags such as `[laugh]`, `[curious]`, `[pause]`, or any
engine-specific markup into `final_response`, `source_answer`, clarification
text, deflection text, compactions, recall text, QRA text, or canonical Memory
answer content. Tags in Memory text corrupt byte spans, hashes, citations,
semantic embeddings, compaction provenance, and future recall.
The routing/final-response split is strict:
```text
/intent
-> returns route/query metadata plus delivery_context only
-> does not return final text
-> must not return delivery_plan because final-text byte spans do not exist
/answer, /clarify, /deflect
-> return clean text plus delivery_plan
-> delivery_plan spans and text_sha256 are computed over the clean text
-> Chatterbox/SPARTA compiles delivery_plan into runtime render_manifest
```
`delivery_context` is early, turn-scoped metadata. It may include listener
evidence, resolved speaker state, classifier source, affect category,
confidence, and routing influence. It must decay by default and must not become
persona memory unless separately promoted with provenance as a durable user
preference or repeated stable pattern. Situational frustration is not persona
memory by default.
`delivery_plan` is final-text metadata. It must be engine-neutral and valid
against the clean response text hash. It is allowed on `/answer`, `/clarify`,
and `/deflect` only after clean text exists.
Freeze this v1 tone vocabulary:
```text
neutral
calm
warm
careful
firm
concerned
curious
deescalating
urgent
```
Freeze these v1 effect keys:
```text
emphasis
pause_before_ms
pause_after_ms
pace_multiplier
intensity_delta
```
Tone influences delivery first, routing second, and never overrides grounding.
A hostile, frustrated, discouraged, or confused turn may make a response calmer,
firmer, more careful, more concerned, or more clarifying. It must not convert an
answerable grounded query into `/deflect` unless safety, off-topic, no-match, or
answerability rules independently justify that route.
SPARTA/Chatterbox may write a runtime `render_manifest` and injectable tags at
playback time. Memory may store or reference playback audit artifacts only as
external evidence. A render manifest is not canonical Memory answer content.
Live proof for this contract in the memory repo:
```bash
cd ${HOME}/workspace/experiments/memory
./scripts/prove-delivery-plan-contract.sh --base-url http://127.0.0.1:8601
```
The proof must show `/intent` has `delivery_context` and no `delivery_plan`,
while `/answer`, `/clarify`, and `/deflect` have clean text plus
`delivery_plan`, with no `emotion_tags`, `chatterbox_tags`, or `voice_policy`
leaking into canonical Memory output.
#### POST /speaker/resolve -- Voice Speaker Identity Product
Use `/speaker/resolve` before `/intent` for voice turns where the listener has
speaker verification or diarization evidence. This endpoint consumes upstream
speaker candidates and returns a memory-safe identity decision:
| Status | Caller behavior |
| --- | --- |
| `known` | Pass `speaker_resolution` to `/intent`, then use `/recall` with returned tags such as `speaker:horus_lupercal`, `user:horus_lupercal`, `persona:horus_lupercal`, and `persona:embry`. |
| `unknown` | Do not run personal memory/QRA recall. Ask the returned identity prompt, for example "Who am I speaking with?" |
| `ambiguous` | Do not choose between profiles. Ask the returned identity prompt or a disambiguating follow-up. |
`/speaker/resolve` is not an audio model. It does not compute embeddings,
transcribe speech, or inspect raw audio. RealtimeSTT, ECAPA, pyannote, or the
listener service owns audio evidence. Memory owns the identity/profile decision
Auf GitHub ansehen