- name
- verify-memory-integrity
- description
- Verify agent memory reachability and budget without modifying anything — orphan detection, broken links, dual-cap usage against the 200-line / 25KB truncation limits, and catalog conformance. Use at the start of a session before trusting memory contents, before and after any index compaction, after a rename or repository move, when a memory directory has grown past a few dozen topic files, or as a periodic health check.
- license
- MIT
- allowed-tools
- Read Bash Grep Glob
- metadata
- {"author":"Philipp Thoss","version":"1.0","domain":"general","complexity":"intermediate","language":"multi","tags":"memory, claude-code, verification, reachability, read-only, maintenance","locale":"de","source_locale":"en","source_commit":"20fe5e96f423ff782f4a492f3310876042e09437","fence_basis_commit":"20fe5e96f423ff782f4a492f3310876042e09437","translator":"(untranslated stub)","translation_date":"2026-08-23"}
# Verify Memory Integrity
Read a Claude Code memory store and report whether it is reachable and within budget — without
changing a byte of it. A store-identity resolution and eight checks cover the two truncation caps,
orphaned topic files, dangling and degraded references, cross-reference resolution, catalog
conformance, topic-file size, and recoverability.
`manage-memory` and `prune-agent-memory` both mutate the store. This skill is the instrument they
are measured with: run it before either one for a baseline, and again afterward to prove the
mutation stranded nothing. An instrument that edits its subject cannot measure a change.
## When to Use
- At the start of a session, before trusting anything the memory index claims
- Before and after an index compaction — to record what was reachable while it still was, then to
prove nothing was stranded. Compaction is the operation the caps make mandatory, and it is also
the operation that breaks reachability
- After a repository rename or move, which invalidates paths inside memory entries
- When a memory directory has grown past a few dozen topic files and nobody has audited the links
- As a periodic health check, on the same cadence as `prune-agent-memory`
**Do NOT use** to fix anything. This skill never writes to the store. Hand its report to
`manage-memory` (to relink or extract) or `prune-agent-memory` (to delete) and re-run it afterward.
## Inputs
| Parameter | Type | Required | Description |
|---|---|---|---|
| `memory_dir` | path | Yes | The store to verify, typically `~/.claude/projects/<project>/memory/`. Never hardcode it: `autoMemoryDirectory` in any settings scope and `CLAUDE_CODE_PROJECT_DIR_NAME` both relocate it |
| `report_path` | path | No | Where the report is written (default: a path under `$TMPDIR`). Must be outside `memory_dir` |
| `topic_max_bytes` | number | No | Per-project topic-file size threshold for check 7. No default — see that step |
| `name_prefixes` | list | No | Filename prefixes that check 5 may strip when resolving (default: `feedback-`, `project-`, `reference-` — hyphen form, because the block normalizes `_` to `-` before it strips) |
## Read-Only Contract
**No `Write`, no `Edit`, no `rm`, no `mv`, no redirect into the store.** `allowed-tools` omits the
mutating tools deliberately and no step below writes inside `memory_dir`, which is what makes the
skill safe every session and its output usable as a before/after baseline.
Each check appends verdict lines to one report file **outside** the store, and the run exits
non-zero if any line begins `FAIL`. Steps run as separate tool calls, so shell state does not
persist: re-export the report path and the store path in each call rather than relying on a trap or
an earlier assignment. The report skeleton, its exit-code rules and the fail-closed guard on a
missing report are in [references/EXAMPLES.md](references/EXAMPLES.md).
## Procedure
### Step 0: Resolve the store before measuring it
Every number below reads whichever directory this step names, so name it first. The
project-directory slug transformation changed between 2026-03-22 and 2026-04-16: an underscore in a
project path used to be preserved and is now converted to a hyphen. So memory written before the
change lives under a slug the harness will never open again, and two project paths differing only by
`_` versus `-` now collide onto one store.
```bash
# Report the RESOLVED path, plus any sibling store differing only in _ vs -.
DIR=<memory-dir> # .../projects/<slug>/memory
STORE=$(cd "$DIR" && pwd -P) && echo "STORE $STORE"
SLUG=$(basename "$(dirname "$STORE")"); ROOT=$(dirname "$(dirname "$STORE")")
KEY=$(printf '%s' "$SLUG" | tr '_' '-') # both spellings collapse to one key
for d in "$ROOT"/*/; do
alt=$(basename "$d")
[ "$alt" = "$SLUG" ] || [ "$(printf '%s' "$alt" | tr '_' '-')" != "$KEY" ] \
|| [ ! -d "$ROOT/$alt/memory" ] \
|| echo "SIBLING $ROOT/$alt/memory — a second store differing only in _ vs -"
done
```
Scan and normalize; do not substitute. Building the candidate with `${SLUG//-/_}` looks equivalent
and silently kills the direction that matters — measured on two real sibling pairs, and explained in
[references/EXAMPLES.md](references/EXAMPLES.md).
**Expected:** One `STORE` line naming the absolute path actually measured, and no `SIBLING` line.
**On failure:** Append `FAIL store identity` and stop before check 1. Unreachable is
indistinguishable from empty, so a budget and an orphan count are both perfectly accurate readings
of the wrong store. Determine which store the session reads, then re-run from here.
### Step 1: Measure the dual-cap budget
`MEMORY.md` is truncated on load at **200 lines or 25KB, whichever comes first** (per the Claude
Code memory documentation). Measure both and act on the larger fraction:
`usage = max(size / 25000, lines / 200)` — warn at `0.80`, rewrite target `0.70`.
Above ~124 characters of content per line the **size cap binds first** — 125 units once the line
separator is counted, since 200 lines carry 199 separators. At the ~150-character entry this skill
targets — a derivation here, not a documented recommendation — the real budget is ~166 lines, not
200. Never report a line count alone, and always name which cap binds.
**Before quoting any figure above**, read [what is documented and what is
derived](references/EXAMPLES.md#what-is-documented-and-what-is-derived) and [what the size cap counts, and on which versions](references/EXAMPLES.md#what-the-size-cap-counts).
**What the size cap counts (measured, not documented).** UTF-16 code units — JavaScript
`String.length` — not UTF-8 bytes and not code points. Inside the BMP a character count is exact; a
byte count over-reports (up to 3x on CJK) and a character count *under*-reports on astral
characters, where one character costs two units.
```python
size = sum(2 if ord(c) > 0xFFFF else 1 for c in text) # UTF-16 code units
```
**Three further measured properties** (Claude Code 2.1.237–2.1.241, linux-x64/WSL2, tool use
disabled; derivation in [references/EXAMPLES.md](references/EXAMPLES.md)): truncation is whole-line,
so `first line dropped` names a line dropped entirely rather than cut mid-way; carriage returns
count, so a CRLF index has a smaller line budget than the same content with LF; and being past the
crossover says only which cap will bite first, not that anything is being truncated yet. State the
boundary as measured across that range, never as a constant.
```bash
# Dual-cap budget. Reads bytes and decodes ONCE: Python's text mode deletes CR
# characters, which silently changes the string the loader actually measures.
IDX=<memory-dir>/MEMORY.md
python3 - "$IDX" <<'PY'
import re, sys
raw = open(sys.argv[1], 'rb').read()
full = raw.decode('utf-8', 'replace')
# Only the content that LOADS counts toward either limit. YAML frontmatter and
# block-level HTML comments are stripped before the index is loaded and are
# excluded from the measurement; a comment INSIDE a fenced code block is not —
# it is preserved and counted (measured, see the skill text). Stripping it too
# under-reports, which hides a truncation that is already happening.
text = re.sub(r'\A---\r?\n.*?\r?\n---[ \t]*\r?\n', '', full, flags=re.S)
kept, fence, cmt = [], False, False
for ln in text.split('\n'):
# An OPEN comment wins over the fence rule (#734): ``` inside one is content,
# not a delimiter, and fence-first let `cmt` outlive it and strip real lines.
if cmt:
cmt = '-->' not in ln
continue
if ln.lstrip().startswith('```'):
fence = not fence
elif not fence and ln.lstrip().startswith('<!--'):
cmt = '-->' not in ln
continue
kept.append(ln)
text = '\n'.join(kept)
units = lambda s: sum(2 if ord(c) > 0xFFFF else 1 for c in s) # UTF-16 code units
lines = text.count('\n') + (0 if text.endswith('\n') else 1)
size = units(text)
lf, sf = lines / 200, size / 25000
# Track BOTH caps and take the minimum. A loop that tests only the size cap and
# checks the line cap in its else-branch reports every large file as size-bound.
acc = cut_size = 0
for i, line in enumerate(text.split('\n'), 1):
acc += units(line) + 1
if acc > 25000:
cut_size = i
break
cut_line = 201 if lines > 200 else 0
cuts = [c for c in (cut_size, cut_line) if c]
binds = 'size' if cut_size and (not cut_line or cut_size <= cut_line) else 'lines' if cut_line else 'neither'
print(f"lines {lines}/200 = {lf:.1%} size {size}/25000 = {sf:.1%}")
print(f"USAGE {max(lf, sf):.1%} -> " + ("OVER CAP - the tail is dropped on load" if max(lf, sf) >= 1.0
else "COMPACT NOW (target 70%)" if max(lf, sf) >= 0.80 else "OK"))
print(f"binds: {binds}" + (f"; first line dropped: {min(cuts)}" if cuts else ""))
print(f"mean {size / lines:.0f} units/line — the size cap binds first above 125")
print(f"utf-8 bytes {len(raw)}; not loaded, so not counted: {units(full) - size} unit(s)")
print(f"astral chars {sum(1 for c in text if ord(c) > 0xFFFF)}")
PY
```
**Measurement-basis note.** The block measures what LOADS: frontmatter and block-level HTML comments
are stripped, a comment inside a fenced code block is **not**. Getting either half wrong misreads in
a different direction — raw over-reports (one store of four here read 70.4% of cap raw against 68.7%
loaded), stripping fenced comments under-reports and hides a truncation already happening. Both,
plus four CommonMark divergences each labelled with the direction it errs: [references/EXAMPLES.md](references/EXAMPLES.md#what-the-strip-does-and-does-not-remove).
**Expected:** Both fractions with both denominators, a `binds:` verdict naming which cap would cut
first (`neither` while both still have headroom), the mean units per line against the crossover,
and `USAGE` under 80%.
**On failure:** At `USAGE >= 0.80` append `FAIL dual-cap` and hand the store to `manage-memory` with
a rewrite target of 70% of the *binding* cap. At `USAGE >= 1.0` the tail is already invisible on
load — record `first line dropped` before anyone edits the file, because the next write moves it.
### Step 2: Detect orphaned topic files
Only the index is loaded automatically, so a topic file nothing links to is unreachable — and
nothing about its own contents will ever reveal that.
```bash
# Reachability: the index is the only file loaded automatically, so a topic file
# that nothing links to is not deprioritized — it is invisible.
DIR=<memory-dir>
python3 - "$DIR" <<'PY'
import os, re, sys
d = sys.argv[1]
text = open(os.path.join(d, 'MEMORY.md'), 'rb').read().decode('utf-8', 'replace')
# HTML comments are stripped before the index reaches the model, and the
# stripped content is excluded from the load limits: a note left in one is
# invisible to the reader, and buys nothing by being cheap.
text = re.sub(r'<!--.*?-->', '', text, flags=re.S)
EXAMPLES = {'file.md', 'example.md', 'topic-name.md'} # format-documentation targets
linked = {os.path.basename(m) for m in re.findall(r'\]\(([^)#\s]+\.md)', text)} - EXAMPLES
on_disk = {f for f in os.listdir(d) if f.endswith('.md') and f != 'MEMORY.md'}
orphans, dangling = sorted(on_disk - linked), sorted(linked - on_disk)
size = lambda names: sum(os.path.getsize(os.path.join(d, n)) for n in names)
tot = size(on_disk) or 1
print(f"topic files {len(on_disk)}; linked {len(linked & on_disk)}")
print(f"ORPHANS {len(orphans)} = {len(orphans)/max(len(on_disk),1):.1%} of files, {size(orphans)/tot:.1%} of bytes")
print(f"DANGLING {len(dangling)} (linked, absent on disk)")
for n in orphans: print(f" orphan {n}")
for n in dangling: print(f" dangling {n}")
PY
```
Five rules this block honors, each learned from a real miss. They govern how its output is read as
much as how it is written, and steps 3, 4 and 6 each pick one up:
1. **Exclude template/example link targets** — a format-documentation line such as
`- [Title](file.md) — hook` otherwise reports as dangling forever, and a check that cries wolf
trains its operator to ignore it.
2. **A prose mention is not reachability.** Require an exact filename match on a real link target.
Report near-matches separately as *degraded references*; never count them as reachable.
3. **HTML comments in the index are not a mitigation.** They are stripped before the index reaches
the model, so a curator note in one is written into a void. Being stripped, it is also excluded
from both load limits — it is free, and unread, which is the worst combination for a note whose
whole purpose is to be seen. A note that must survive has to be a plain markdown line. (Before
v2.1.211 the raw file was measured, so such a note did also consume budget.)
4. **Parse frontmatter, do not grep it.** A `^type:` regex misses a field nested under `metadata:`
and reports a conformance failure that does not exist.
5. **Report both denominators, labeled.** File share and byte share are different numbers, never
interchangeable — a store can be 40% orphaned by file count and 5% by bytes.
**Expected:** `ORPHANS 0`, with both denominators printed and labeled. `topic files` equals `linked`.
**On failure:** Append `FAIL orphans` with the file list and both percentages. Do not delete — an
orphan is as likely to be a valuable file whose index line a compaction dropped. `manage-memory`
relinks; `prune-agent-memory` decides deletion.
### Step 3: Resolve dangling links
Step 2's `DANGLING` list is the sole verdict source for broken links — never add a second extractor
that can disagree with it. This step adds the reading: every entry is either a real break or a
format-documentation target belonging in the `EXAMPLES` set (rule 1).
```bash
# DIAGNOSTIC ONLY — Step 2 holds the verdict. Lists in-scope suspects: every
# target that is not a flat `*.md`. Full link-form census in EXAMPLES.md.
grep -oE '\]\([^)]+\)' <memory-dir>/MEMORY.md | tr -d ']()' \
| grep -vE '^[^/]+\.md$|^#|://' \
|| echo "LINK FORMS all flat *.md — no sub-path, anchor or URL targets"
```
Anchor-only and URL targets are out of scope. A sub-path target is in scope and a likely break: the
store is flat, so a link carrying a directory component usually survived a move.
**Expected:** `DANGLING 0` after every format-documentation target has been accounted for by the
`EXAMPLES` set, and no `sub-path` targets.
**On failure:** Append `FAIL dangling` with the list. For a documentation example, extend `EXAMPLES`
in the Step 2 block rather than suppressing the check: an exclusion list is auditable, a disabled
check is not.
### Step 4: Report degraded references
A filename that appears in the index only as prose or inline code is not a weak link, it is not a
link (rule 2). It reads as reachable to a human scanning the index and is invisible to the mechanism
that makes files reachable. Counted separately, never as reachable — the files are already in Step
2's orphan list, and this step explains why they looked fine.
```bash
# Degraded references: named in the index, never as a link target.
python3 - <memory-dir> <<'PY'
import os, re, sys
d = sys.argv[1]
text = open(os.path.join(d, 'MEMORY.md'), encoding='utf-8', errors='replace').read()
text = re.sub(r'<!--.*?-->', '', text, flags=re.S)
linked = {os.path.basename(m) for m in re.findall(r'\]\(([^)#\s]+\.md)', text)}
GitHubで見る