| name | auto-learn |
| description | | |
A wrong or bloated learning poisons every future run — this skill's failure mode is worse than
learning nothing. Non-negotiable:
1. LEARN FROM ARTIFACTS, NOT VIBES. Candidates come from the run's structured evidence — the
checkpoint's register/openItems, blocked units with their notes, guard-block reasons, fix-loop
iteration counts, delegation degradations, verifier rejection ledgers — each candidate carries its
evidence reference. No evidence, no learning.
2. VERIFY BEFORE WRITING. Every candidate passes `adversarial-verify`: is it TRUE (matches the
evidence), GENERAL (a pattern, not a one-off flake), and ACTIONABLE (changes a future decision)?
One flaky run does not make a law. Rejected candidates are dropped, not hoarded.
3. WRITE TO WHERE CLAUDE CODE ACTUALLY LOADS CONTEXT — never a bespoke file nothing reads. A learning
only changes behavior if it lands in native, auto-loaded memory: **CLAUDE.md** (loaded in full every
session), a path-scoped **`.claude/rules/.md`** (loaded when the agent touches that area), or
**auto memory** (`~/.claude/projects//memory/`, loaded every session). Machine defects (a
skill/guard/template gap) → surfaced to the user, never self-suppressed. NEVER invent a `.ulpi/`
learnings file — nothing loads it, so it is a lesson written into the void.
4. DEDUPE AND CURATE — memory is a working set, not an archive: merge with an existing entry instead of
appending a twin, prune entries a later run superseded or disproved, convert relative dates to
absolute, and keep CLAUDE.md within its line budget. Cap: at most 5 new learnings per run.
5. NEVER learn secrets/tokens/URLs-with-credentials, and never rewrite human-written memory — the
same stamps-and-preservation rules as auto-map.
Auto Learn
Inputs
$run: a checkpoint path or run id to harvest; defaults to the most recent .ulpi/runs/*.json.
Overview
The difference between an autonomous machine and a self-improving one is whether run N+1 knows what
run N paid to find out. Every blocked task, guard block, thrash loop, and rejected finding is evidence
about how THIS repo and THIS process actually behave. This skill turns that evidence into verified,
deduplicated learnings and writes them into the memory Claude Code already loads every session —
CLAUDE.md, .claude/rules, auto memory — so the next run starts with the lessons in context
automatically, without any phase having to remember to read a file. That is what closes the loop.
Phase 0: Harvest the run's evidence — it is CODE
node <skill-dir>/scripts/harvest-run.mjs <checkpoint.json> --json
The harvester extracts candidates mechanically from the checkpoint: blocked/dep_blocked units with
their notes, gate failures from the register, delegation degradations, units that needed many fix
iterations, and openItems by phase. Supplement with what the checkpoint can't hold: guard-block
messages you observed, verifier rejection patterns (repeated false-positive shapes), and anything the
user corrected mid-run. Every candidate keeps its evidence pointer.
Success criteria: a candidate list where each entry cites its evidence; nothing from memory alone.
Phase 1: Verify and generalize (the anti-poison gate)
For each candidate, adversarial-verify with three questions — TRUE (does the evidence actually
support it — reread the artifact, don't trust the summary)? GENERAL (would it recur — or was it a
one-off flake/typo)? ACTIONABLE (what future decision changes)? Then rewrite survivors into the
canonical learning shape:
- <the rule> — WHY: <one line, from evidence> — APPLY: <what to do differently> (evidence: <ref>, <date>)
Success criteria: ≤5 survivors, each true, general, actionable, and evidence-linked.
Phase 2: Route each learning to native, auto-loaded memory
Claude Code loads these every session (or on matching-file access) with no prompting. Put each learning
in the one that will actually reach the next run:
| Learning is about… | Route to (Claude Code auto-loads it) | How |
|---|
| A project-wide convention, invariant, or plan-shape rule the run kept violating (tasks too fat here, a validate form that lies, a boundary that bit) | CLAUDE.md — loaded in full every session | append to a stamped ## Learnings section, or hand it to auto-map (which owns CLAUDE.md) |
| A lesson that only applies to one area (an API rule, a DB/migration gotcha, a test-suite quirk) | .claude/rules/<area>.md with paths: frontmatter — loads when the agent touches those files | create or append the rule; scope its paths: glob |
| An environment or tooling quirk the run discovered (build flake, slow suite, a service dependency, a CI oddity) | auto memory (~/.claude/projects/<project>/memory/) — MEMORY.md loaded every session | a topic file plus one MEMORY.md index line |
| A defect in the machine itself (a skill gap, a guard bypass, a template bug) | the user — reported, never self-patched | surface it in the Output Contract |
| A cross-agent lesson both Codex AND Claude Code must load (the shared native context file) | the nearest AGENTS.md — read by Codex and Claude Code alike | route it through scripts/route-learnings.mjs (below) — DRY-RUN first, then --apply |
The AGENTS.md router (DRY-RUN-FIRST, fail-closed) — Codex + Claude Code
AGENTS.md is the ONE native memory file both Codex and Claude Code read, so a lesson that must reach
either agent belongs there. Because that file is shared, auto-loaded, and cross-agent, writing to it by
hand is how a bad lesson silently poisons every future run. Route it through code instead:
node <skill-dir>/scripts/route-learnings.mjs learnings.json --root <project-root>
node <skill-dir>/scripts/route-learnings.mjs learnings.json --root <project-root> --apply
Each learning in the input ({"learnings":[…]} or a bare array) MUST carry five fields — id (stable),
rule (actionable), evidence (a real artifact ref), verification (an affirmative marker), and
scope (a repo directory the lesson applies to). The router's fail-closed contract:
- Dry run is the default. No
--apply, no writes — you get a patch manifest (per-target before/after
SHAs, the survivors, and every rejection with its reason) to inspect first.
--apply edits ONLY the stamped, auto-learn-owned block of the target AGENTS.md
(<!-- BEGIN auto-learn:learnings … --> … <!-- END … -->). Every other byte of that file is
preserved; prior generated learnings are merged forward (dup ids merge evidence), never dropped.
- It targets the nearest
AGENTS.md at or above the learning's scope, falling back to the root
AGENTS.md. It writes ONLY files named AGENTS.md inside the project root — it NEVER edits CLAUDE.md
or private Codex memory.
- ≤5 survivors per run; duplicate ids merge evidence. More than five distinct additions ⇒ the whole
run refuses to mutate anything (
refused: "too_many_additions").
- No mutation on any of: missing/blank required field, placeholder/absent evidence, a non-affirmative
(
unverified/…) verification, a secret-shaped value anywhere in the record, a scope that path-
traverses (../absolute/escapes root), a nonexistent scope directory, or a machine/environment
defect (kind:"machine_defect", or defect-shaped rule/evidence text). Defects are reported for the
user, never self-patched into shared memory.
Use this router for AGENTS.md; use auto-map for CLAUDE.md/.claude/rules and auto-memory for the
Claude-only layers. Dedupe on write: if an existing entry covers it, strengthen that entry (add the new evidence ref)
instead of appending a twin; prune superseded/disproved entries while you are there. Keep CLAUDE.md and
each rules file within budget — they load into every session, so bloat taxes all future work.
Success criteria: every survivor landed in exactly one native-loaded location; nothing written to a
file Claude Code does not load; memory deduped and within budget.
Phase 3: Close the loop — the feed-forward is now AUTOMATIC
Because the learnings live in native, auto-loaded memory, the next run picks them up with no extra step —
with one nuance to state honestly: CLAUDE.md and auto-memory MEMORY.md load in full EVERY session, so
lessons routed there close the loop unconditionally; a .claude/rules/<area>.md lesson loads only when a
future session touches files matching its paths: glob (path-scoped by design — that is the point of an
area-specific rule). So route a lesson the next run must ALWAYS see to CLAUDE.md / auto-memory, and an
area-specific one to .claude/rules. Confirm and report that each learning actually landed in a loaded
location (/memory lists what is loaded; check the CLAUDE.md ## Learnings section and any new
.claude/rules file). In pipeline composition this skill runs BEFORE auto-map, which then verifies the
CLAUDE.md/rules the learnings touched against the real repo.
Success criteria: every lesson is in a location Claude Code confirms it loads; none stranded in a
file nothing reads.
Common Rationalizations
| Rationalization | Reality |
|---|
| "Obvious lesson, no need to verify." | Obvious-feeling lessons from one bad run are how false laws get written. Verify against the evidence or drop it. |
| "Keep every learning — more is better." | CLAUDE.md loads into every session. Bloat taxes every run; 5 curated beats 50 hoarded. |
"Drop it in a .ulpi/learnings.md and have the skills read it." | Claude Code doesn't load that file. It's a lesson written into the void. Use CLAUDE.md, .claude/rules, or auto memory — what actually loads. |
| "Just append; dedup later." | Later never comes. Twins drift apart and contradict; merge on write. |
| "That failure was the skill's fault — tweak my own instructions quietly." | Machine defects go to the user. Self-patching hides systemic bugs from the person who owns the system. |
| "Record it in the transcript summary." | Transcripts die with the session. Only routed, durable layers change the next run. |
| "One flake = a rule." | One occurrence is an anecdote. Two with the same evidence shape is a pattern. Generalize only what recurs or is structurally certain. |
Red Flags
- A learning with no evidence pointer, or one contradicting the artifact it cites.
- A learning written to any file Claude Code does not auto-load (a
.ulpi/ note, a random doc): it will
never reach the next run.
- CLAUDE.md or a rules file growing past its budget or accumulating near-duplicate entries.
- A machine defect recorded as a repo note instead of surfaced to the user.
- The same mistake appearing in two consecutive runs' registers — the loop is NOT closing; escalate.
Guardrails
- Never write an unverified or evidence-free learning; never learn from memory alone.
- Never write a learning to a file Claude Code does not load (CLAUDE.md /
.claude/rules / auto memory
only); never invent a .ulpi/learnings.md.
- Never exceed 5 new learnings per run; keep CLAUDE.md and rules files within budget.
- Never duplicate — merge and strengthen; prune superseded entries on every pass.
- Never store secrets; never touch human-written memory outside stamped sections.
- Never self-patch machine defects silently — they belong to the user.
When To Load References
scripts/harvest-run.mjs — mechanical candidate extraction from a checkpoint (Phase 0). CI-tested by
scripts/test-harvest.sh (every signal class extracted with evidence citations; unreadable → exit 2).
scripts/route-learnings.mjs — the DRY-RUN-FIRST, fail-closed router that writes verified learnings to
the shared cross-agent AGENTS.md (Phase 2, Codex + Claude Code). CI-tested by
scripts/test-learn-route.sh (dry-run default, --apply byte-preservation, the ≤5 cap + dup-merge,
and every no-mutation refusal). Use it for AGENTS.md; use auto-map for CLAUDE.md/.claude/rules.
adversarial-verify (skill) — the TRUE/GENERAL/ACTIONABLE gate (Phase 1).
auto-map (skill) — the routing target for repo-fact learnings (Phase 2).
checkpoint-resume (skill) — the artifact format this skill harvests.
Output Contract
Report:
- run harvested (checkpoint path) + candidates found by the harvester vs. added manually
- verification: survivors vs. rejected (with the rejection reason)
- routing: which learning landed where (CLAUDE.md /
.claude/rules/<area>.md / auto memory / user-report)
- machine defects surfaced to the user (never self-patched)
- confirmation each lesson is in a location Claude Code auto-loads (entries merged/pruned, within budget)