spec
Create a technical specification — `/spec short` for single-pass plans, `/spec` for full team-based investigation
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Create a technical specification — `/spec short` for single-pass plans, `/spec` for full team-based investigation
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
SOC 職業分類に基づく
Drive a feature end-to-end across multiple protocol sessions — the coordinator role's protocol home
Create a GitHub pull request from the current branch, deriving the PR body from the associated work item's plan and notes.
Holistic multi-lens PR review with adaptive lens selection, cross-lens synthesis, and structured findings; --self runs the same pipeline as an author's self-review. Use individual lens skills (/pr-correctness, /pr-security, etc.) for focused single-concern analysis.
Check project status, remaining tasks, and session context — USE FIRST when asked 'what's left', 'what should I do', 'remaining work', or status questions. Also: create, update, archive, search work items.
Focused lens review: trace the impact of PR changes on code outside the diff. Use /pr-review for integrated multi-lens coverage.
Focused lens review: trace logic paths for correctness bugs in a PR. Use /pr-review for integrated multi-lens coverage.
| name | spec |
| description | Create a technical specification — `/spec short` for single-pass plans, `/spec` for full team-based investigation |
| user_invocable | true |
| argument_description | [short] [--yes] [--model <id>] [name or description] — existing work item name, or a freeform description to start from. `--model` overrides every per-role binding (lead/researcher/advisor) for this invocation only; otherwise per-role models come from `resolve_model_for_role`. |
Produces a plan.md inside a work item's _work/<slug>/ directory.
/spec short)Single-agent path: the spec-lead reads key files directly (Step 2 --short branch) and drafts the plan without dispatching a researcher team. Single-context work is the sanctioned norm, not a degraded mode — one context that reads the code itself serves most specs well.
The --short conditional activates at Step 2 only. From Step 3 onward, short and full paths share every step: collect findings and emit Tier-2 artifacts, strategy gate, synthesis, design ceremony, task review, post-research extraction, post-plan ceremony, and terminal finalization.
/spec)Team-based divide-and-conquer: the spec-lead composes an investigation plan table (Step 2 full branch), dispatches parallel researcher agents, collects findings, emits Tier-2 artifacts, and then synthesizes — following the shared downstream steps. Team ceremony engages where scale demands it: investigation breadth one context cannot hold, or unknowns that genuinely parallelize.
Sequencing constraint: Do not dispatch research agents before completing Step 2. The investigation plan is a completeness checklist and user approval gate, not just a dispatch list.
The verbs prepare evidence and persist decisions; they do not make them. Keep these kernels in lead prose: investigation questions, complexity labels, strict/permissive applicability, synthesis, contradiction decisions, phase/task decomposition, evaluator normalization, whether feedback changes the plan, and every harness-native dispatch or teardown call. If a verb appears able to infer one of these, stop at the boundary and supply the missing judgment explicitly.
Parse arguments:
TRACK=short when the first arg after /spec is short; otherwise set TRACK=full. The investigation step (Step 2) follows that declared track.--yes is present, skip all interactive confirmation gates (auto-proceed through investigation plan confirmation, strategy gate, confirm understanding, and task review).--model <id> is present, set MODEL_OVERRIDE=<id> and export the per-role overrides for this invocation: export LORE_MODEL_LEAD=<id> LORE_MODEL_RESEARCHER=<id> LORE_MODEL_ADVISOR=<id>. Otherwise leave MODEL_OVERRIDE empty and let ceremony-scoped role resolution choose each model. An explicit flag without a value is an error, not a request for the configured default.Resolve the knowledge path and standing defaults:
lore resolve
lore defaults
Set KNOWLEDGE_DIR to the first result and WORK_DIR to $KNOWLEDGE_DIR/_work. The second renders the standing defaults in force (settings-derived role/model maps, ceremony registrations, sampling rates, preference directives cited by title); treat its output as binding for this run.
Run the read-only startup verb. It owns work-reference resolution, plan-state classification, active-framework/model resolution, template-version stamping, and source provenance. It never creates a work item, repairs an index, or chooses a protocol route:
START_INPUT=${INPUT:-__branch_inference__}
START_ARGS=("$START_INPUT" --branch "$(git rev-parse --abbrev-ref HEAD 2>/dev/null || true)" --json)
[[ "$TRACK" == "short" ]] && START_ARGS+=(--short)
[[ -n "$MODEL_OVERRIDE" ]] && START_ARGS+=(--model "$MODEL_OVERRIDE")
START=$(lore spec start "${START_ARGS[@]}")
Required version-1 fields are schema_version, resolved, slug, archived, plan_state, intent_anchor, strategy_present, active_framework, effective_lead_model, track, lead_template_version, and provenance. Missing declarations and unknown flags are errors; there is no default route for malformed input. Exit 2 means ambiguity: ask the user to select from the ordered candidates, then re-run with the exact slug.
Choose the route from the returned facts — this is lead judgment, not verb output:
resolved=false with user input: treat the remaining input as a freeform description and continue to Goal Refinement.resolved=false without user input: ask what the user wants to spec; never turn the branch-inference sentinel into a work item.archived=true: warn and wait for explicit confirmation.plan_state=synthesis-complete: load plan.md, read any ## Strategy silently, and continue at Step 5.1.plan_state=investigations-only: load the persisted findings and any strategy, then continue at Step 5.plan_state=follow-up-needed: read the open questions and design targeted follow-up investigations.plan_state=incomplete: present the persisted material for discussion; offer the strategy gate before synthesis when no strategy exists.plan_state=none: continue to Step 2.When intent_anchor is non-null, preserve it verbatim — downstream gates (the Step 5.6 verifier, /implement anchor prompts) detect drift by string comparison, so a paraphrase breaks the audit chain even when intent is preserved. It is the user-visible capability boundary, not a suggestion the spec may silently narrow. Startup reports the anchor; only the lead judges whether the emerging design still covers it.
When the input is a freeform description rather than an existing work item:
AskUserQuestion. Target scope boundaries, constraints, and approach preferences. Do NOT ask questions answerable by reading the codebase./work create flow and pass --intent-anchor with an interpreted one-sentence capability statement that names the user-visible outcome. Audit it for looseness — alternatives ("X or Y" admits doing just one), comparatives without targets ("better," "faster" with no acceptance bar), vague verbs ("support," "improve" without naming the bar), and the meta-instruction degenerate case (user said "make a work item for that," no capability named) — before committing. The clarifying questions in step 2 should have closed most looseness; if any remains, ask one targeted question via AskUserQuestion before creating. See /work create intent-anchor guidance and [[knowledge:conventions/protocol/work-item-intake-should-store-neutral-intent-ancho]].If the user's description is already specific enough (clear scope, stated constraints, obvious approach), skip to step 4 — don't ask questions for the sake of asking.
--short--short flag present)From conversation context and the work item, identify 3-8 key files to read.
Search the knowledge store: lore search "<topic>" --type knowledge --scale-set subsystem,implementation --json --limit 5. Read relevant entries.
Check the knowledge store index for relevant domain files.
Read the files yourself — do NOT spawn subagents.
Note key findings as you go.
notes.md or other uncommitted prose (not committed code or a commons entry), verify the claim against HEAD before adopting it. Notes age faster than code; a stale note silently shapes scope._meta.json) or inserts a script into an existing control-flow gate, trace three seams before drafting: (a) does an existing gate/guard intercept the new state? (b) does the denormalized read model (e.g., _index.json) project the field, or is it invisible to consumers? (c) does ordering (archive/move) change where a later step finds the file?Prepare discovery evidence without assigning meaning:
DISCOVERY=$(lore spec discover "$SLUG" --json)
The version-1 result contains coverage, candidates, and provenance. Coverage names every scanned, missing, or unreadable source stratum; candidates retain each source's own rank and score. The verb excludes canonical Lore skill/agent identities structurally, but it never combines rankings or emits matched, binding, or applicability fields.
The BM25 strata key on a single query seed — the work-item title, unless repeatable --seed <token> flags replace it (tokens join into one query; the wholesale tree scans are seed-insensitive and enumerate everything regardless). The seed is the only query-driven input, which is what lets /implement's close re-run this same enumerator with seeds derived from the shipped diff: the two passes miss independently, and a norm the title-seeded pass ranks below the cutoff gets a second chance at close. A descriptive work-item title is what makes the spec-time half of that pair earn its keep.
Apply the two discovery judgments yourself:
**External skill discovery:** with the considered set and the matched set; Matched: none is valid.**Preference and convention discovery:** with coverage counts and the surfaced backlinks; Surfaced: none is valid.Candidate enumeration is hands work; strict/permissive applicability is head work. Do not ask the verb to collapse that boundary.
Present a context summary and offer the strategy gate (Step 4 below). If --yes, skip the strategy prompt.
--short flag)From the feature description, identify 3-7 focused investigation questions. Each should target a specific codebase concern, be answerable by exploring files, and be independent enough to run in parallel.
Always include two mandatory fixed investigations (both count toward the 3-7 total): a. External skill and agent applicability (strict)... b. Preferences and conventions applicability (permissive)...
a. External skill and agent applicability (strict) — which installed external (non-lore) skills and agent templates should be invoked during implementation of this work item. Lore-managed skills (/spec, /implement, /work, /memory, /remember, /retro, /evolve, /renormalize, /bootstrap, /pr-*, /codex-*) are excluded — protocol toolchain, not advisors. The researcher filters them out before reporting matches. Key files: <skills_dir>/*/SKILL.md, <agents_dir>/*.md (resolve via resolve_harness_install_path skills / resolve_harness_install_path agents); exclusion list comes from the canonical Lore source repo (source ~/.lore/scripts/lib.sh && printf '%s\n' "$LORE_REPO_DIR" + /skills/ and /agents/). Do not use resolve-repo.sh here — it returns the project's knowledge store, not the Lore source tree. Match criterion is strict — include only skills whose stated domain plausibly contributes to this work item's implementation.
b. Preferences and conventions applicability (permissive) — which entries from preferences/, conventions/, and cross-cutting-conventions/ the work might need to honor. Inclusion criterion is permissive — the inverse of skill discovery. The test is "is it possible the work might need this" — not "will we definitely apply it." Err on over-inclusion; synthesis culls. Missing an applicable preference is worse than carrying an inapplicable one through review. Key files: $KDIR/preferences/, $KDIR/conventions/, $KDIR/cross-cutting-conventions/ (full enumeration; absent = zero) plus BM25 from lore search "<topic>" --type knowledge --scale-set subsystem,implementation --limit 10 and --scale-set abstract,architecture --limit 5.
Check the knowledge store index for file hints per investigation.
Assess complexity for each investigation: simple (1-2 files), moderate (3-5 files), complex (6+ files or cross-cutting).
Present the Investigation Plan to the user:
## Investigation Plan
| # | Area / Topic | Key Files | Complexity |
|---|-------------|-----------|------------|
| 1 | External skill and agent applicability — strict, lore toolchain excluded *(mandatory — do not remove)* | `<skills_dir>/*/SKILL.md`, `<agents_dir>/*.md` | simple |
| 2 | Preferences and conventions applicability — permissive *(mandatory — do not remove)* | `$KDIR/preferences/`, `$KDIR/conventions/`, `$KDIR/cross-cutting-conventions/` | simple |
| 3 | <topic> | `file1`, `file2` | simple |
...
Proceed, or adjust?
Wait for user confirmation. If the user requests adjustments, revise and re-present. If --yes, dispatch immediately without confirmation.
Prepare discovery evidence, then make the applicability decisions in lead prose:
DISCOVERY=$(lore spec discover "$SLUG" --json)
Use the returned coverage to see what was scanned or missing. Choose strict external-skill/agent matches and permissive preference/convention candidates yourself. Record matched external skills in $SKILL_INVOCATION_MAP; invoke them inline when a researcher consults that domain. /spec spawns no advisor agents on the default route.
Serialize the approved investigation plan to a temporary JSON file and hold its path as $INVESTIGATIONS_JSON. The version-1 contract is exact — unknown or missing fields refuse, array order is dispatch order, and every prefetch row declares its scale rather than inheriting a default:
{
"schema_version": 1,
"track": "full",
"investigations": [
{"id": "external-skills-agents", "kind": "fixed", "question": "External skill and agent applicability ...", "complexity": "simple", "prefetch": [{"query": "<topic>", "scale_set": ["subsystem", "implementation"]}]},
{"id": "preferences-conventions", "kind": "fixed", "question": "Preferences and conventions applicability ...", "complexity": "simple", "prefetch": [{"query": "<topic>", "scale_set": ["abstract", "architecture"]}]},
{"id": "<lead-authored-id>", "kind": "lead-authored", "question": "<lead-authored question>", "complexity": "simple|moderate|complex", "prefetch": [{"query": "<topic>", "scale_set": ["implementation"]}]}
]
}
Exactly one fixed external-skill/agent question and one fixed preference/convention question are mandatory. The lead owns every question, complexity label, prefetch query, and scale declaration; the verb validates but never invents them.
Open the reusable dispatch artifact:
DISPATCH=$(lore spec open "$SLUG" --investigations "$INVESTIGATIONS_JSON" --json)
open returns created | reused | recovered | replaced, or refuses with the repair target. Its canonical spec-dispatch.json carries input_fingerprint, source_fingerprint, the source manifest, ordered directives, empty lead-side handle slots, and teardown payloads. It never calls a harness tool and never persists live handles.
open also renders and validates lore dispatch guidance before publication or deterministic replay. A rendering failure refuses the whole dispatch wave before any researcher launches, and replay is eligible only while the stable guidance identity remains current. Each directive carries the admitted block in payload.dispatch_guidance, with that identity in the source fingerprint.
Bind ADAPTER, RESEARCHER_MODEL, and RESEARCHER_TEMPLATE_VERSION from the returned source manifest and directives. Do not resolve a second set after publication; that would make the executed payload differ from the artifact being resumed.
Execute the returned directives in ordinal order. For each one, prepend its exact payload.dispatch_guidance verbatim before the researcher template, investigation question, and prior knowledge. Do not rerender or edit the block at launch: all researchers in this deterministic dispatch artifact consume the block admitted by open. A missing block or failed adapter admission ends before the native spawn. This is the harness-dispatch judgment kernel: translate each directive into the active harness's native spawn call, decide when to launch it, and store returned handles only in the lead's in-memory handle map. Populate the matching teardown payload with the live handle at shutdown time; do not write handles back into spec-dispatch.json.
team_messaging != full removes only shared team state. Codex still executes researcher fanout while subagents=partial, then its adapter teardown resolves to lead-mediated TaskUpdate status=completed. Collapse to the short branch only when subagents=none. On a full team-messaging harness, the lead may create the shared team before executing spawn directives and tear it down after researcher completion.
As researcher messages arrive (or after direct file reading in short branch):
Write each finding to the ## Investigations section of plan.md using the investigation entry format from the Plan.md Template below.
Preserve **Findings:** verbatim — copy findings exactly as reported.
Preserve **Observations:** verbatim — copy researcher observations exactly as reported. Do not rephrase, merge, or summarize. These are mechanism-level patterns, design rationale, and structural footprint signals that feed the Step 5.4 capture step.
Emit Tier-2 artifacts — for each researcher assertion (full branch) or lead-observed task-scoped grounding claim (short branch):
claim, file, line_range, exact_snippet, normalized_snippet_hash, falsifier, significance) plus producer/template provenance and change_context (diff_ref, changed_files[], summary). changed_files[] must include the row's file; summary should name why the current investigation/change made the claim relevant. exact_snippet and normalized_snippet_hash are REQUIRED for every row: exact_snippet is the verbatim content at file:line_range that grounds the claim, and normalized_snippet_hash is the sha256 hex of the v1-normalized snippet. Compute the hash via the canonical helper — do NOT inline the recipe:
python3 ~/.lore/scripts/snippet_normalize.py --hash <<<"$SNIPPET"
The v1 normalization recipe (curly→straight quotes, \s+→single space, trim, sha256 lowercase hex) lives only in scripts/snippet_normalize.py. See architecture/artifacts/tier2-evidence-schema.md for the full schema.echo '<json-row>' | bash ~/.lore/scripts/evidence-append.sh --work-item <slug>
evidence-append.sh is the sole writer of $KDIR/_work/<slug>/task-claims.jsonl; it rejects missing snippets and invalid or mismatched normalized hashes. On rejection, fix and retry the row or log the failure to execution-log.md. Never write the JSONL directly — direct writes bypass validation and are treated as corrupt.$KDIR/_work/<slug>/evidence.md. Do not write a mirror entry for a row that failed validation.task-claims.jsonl and evidence.md may be absent — absence means "no Tier-2 claims captured this session," not "work was fully verified."Full branch only: When all investigation tasks are complete:
TaskUpdate task_id=<handle> status=completed directive. Do not infer teardown from team_messaging: the prepared payload and current adapter are the contract.TeamDelete (Claude Code only; opencode/codex adapters require no explicit teardown — runtime owns lifecycle).Append an investigation summary to execution-log.md:
printf 'Investigations: %d\nTopics: %s\n' \
"<N>" "<comma-separated investigation topics>" \
| bash ~/.lore/scripts/write-execution-log.sh --slug <slug> --source spec-lead --template-version "$RESEARCHER_TEMPLATE_VERSION"
Journal the investigation milestone. The findings, Tier-2 rows, and investigation summary are all durable at this point, so a hosted session emits one step_completed row for the parent spec session — individual researcher reports and Tier-2 appends never emit steps:
if [[ -n "${LORE_SESSION_INSTANCE:-}" && -n "${LORE_SESSION_SLUG:-}" && -n "${LORE_SESSION_TYPE:-}" ]]; then
bash ~/.lore/scripts/session-step.sh \
--step-id spec:investigation --step-label "Investigation complete" \
|| echo "[spec] Warning: investigation step not journaled; the persisted artifacts remain authoritative." >&2
fi
The env gate is the hosted-session test — an unhosted run skips silently. Replay is idempotent, and a failed append warns and moves on; it never rolls back the milestone it was reporting.
Before synthesizing, offer the user a chance to shape the plan.
plan.md for a ## Strategy section. If found, read it silently and proceed with it as shaping context — do not re-prompt.## Strategy exists, present a context summary (either compressed investigation summary for full branch, or key-findings summary for short branch) and prompt:
Any strategy to apply to the plan? (Enter to skip)
plan.md immediately:
## Strategy
<user's strategy verbatim>
Then proceed with the strategy as additional shaping context.Always present the prompt — do not skip because scope seems clear. Exception: if --yes was passed, skip this step entirely.
Before synthesizing, check for worker-surfaced concerns:
KDIR=$(lore resolve)
SC_FILE="$KDIR/_work/<slug>/surfaced_concerns.jsonl"
[ -f "$SC_FILE" ] && cat "$SC_FILE"
If present and non-empty, read each pending entry (no status field = unresolved):
## Open Questions## Design Decisions open question or refine the relevant decisionThis step is read-only — do not modify surfaced_concerns.jsonl.
Produce the conceptual frame first before committing to phase breakdown.
Synthesis organizes the itemized findings; it does not narrate over them. Keep each finding's provenance intact and cite items rather than restating them in looser words — prose that absorbs the itemized record is where evidence quietly drops out. The Narrative section is the deliberate exception: it tells the story, while the record underneath stays itemized.
Goal — what we're building/changing and why (1 paragraph). Write it in a controlled register, unconditionally: sentences of at most 25 words, active voice, present tense where possible, one idea per sentence, common words over rare ones, noun clusters of at most three words, no internal vocabulary except technical names grounded at first use. Certified simplified-English compliance is not claimed; the rules bind as written.
Design Decisions — use the ### DN: Title format from the template. Each decision requires **Decision:**, **Rationale:**, **Alternatives considered:**, and **Applies to:** fields. Number decisions sequentially (D1, D2, ...).
Draft Narrative — synthesize goal and chosen approach into a ## Narrative section (1-2 paragraphs). Place it after ## Goal. Write for a reader who wants the story without reading all sections. Draw from Goal and Design Decisions. Omit file paths and task lists. Use the Goal's controlled register where it costs no precision; when the register and technical precision conflict, precision wins — say the precise thing.
Architecture Diagram (conditional) — after drafting Narrative, include a ## Architecture Diagram section when the work touches 2+ distinct modules.
Read diagram conventions:
cat ~/.lore/claude-md/review-protocol/followup-template.md
Diagram types: call chain (invocation paths), state machine (state transitions), data flow (data transforms). Write a plain-text ASCII diagram inside a fenced code block using box-drawing characters. Do NOT use Mermaid or other diagram DSLs — the TUI renderer cannot interpret them.
Consumption-verification checkpoint — before finalizing synthesis, report the outcome for each prefetched commons entry you actually checked against code during investigation. Held and contradicted both count — a confirmation is signal, not ceremony. Skip entries you never tested; grounded-or-nothing means every report needs the code anchor trio, so an entry you can't anchor is an entry you didn't verify:
# Entry confirmed by investigation:
lore verify <knowledge-path> held \
--source spec-lead \
--protocol-slot Synthesis \
--cycle-id "spec-<topic>-$(date +%Y-%m-%d)" \
--template-version "$LEAD_TEMPLATE_VERSION" \
--file <absolute-path> --line-range <N-M> --exact-snippet "<verbatim code>"
# Entry falsified by investigation — additionally lands one pending row in
# $KDIR/_work/<slug>/consumption-contradictions.jsonl; lore audit consumes it as priority-input:
lore verify <knowledge-path> contradicted \
--source spec-lead \
--protocol-slot Synthesis \
--cycle-id "spec-<topic>-$(date +%Y-%m-%d)" \
--template-version "$LEAD_TEMPLATE_VERSION" \
--file <absolute-path> --line-range <N-M> --exact-snippet "<verbatim code>" \
--work-item <slug> \
--rationale "<why the code falsifies the entry>" \
--claim-text "<the entry assertion being contradicted>" \
--falsifier "<what evidence would disprove>"
Events land in $KDIR/_trust/trust-events.jsonl (contract: architecture/trust-ledger/README.md in the knowledge store). Run these from the source repo's root, not from inside the knowledge store — lore resolves the store and records branch provenance from the current directory. Emission is non-blocking — synthesis continues immediately; re-running an identical invocation is a silent no-op (the writers dedupe).
Present the abstract plan (Goal, Design Decisions, Narrative, Architecture Diagram) to the user for review.
Discovery findings integration:
Related skills block (strict): If the discovery researcher (full branch) or Step 2 skill scan (short branch) reported matched external skills, add a **Related skills:** block to the ## Context or ## Investigations section. Lore-toolchain skills are not eligible for this block — they're protocol, not advisors:
**Related skills:**
- /external-skill-name — why this skill is relevant to this work item
Related preferences/conventions block (permissive — audit manifest): If the discovery surfaced any entries from preferences/, conventions/, or cross-cutting-conventions/, add a **Related preferences/conventions:** block to the same section. Include every entry the discovery surfaced under the permissive criterion — workers can dismiss inapplicable ones at implement time; missing applicable ones is the worse failure:
**Related preferences/conventions:**
- [[knowledge:preferences/<entry>]] — what to honor at implement time (1 line)
- [[knowledge:conventions/<entry>]] — what to honor at implement time (1 line)
This block is the audit manifest, not the worker delivery channel. It exists so a reviewer (and the post-plan ceremony) can see every preference/convention discovery surfaced. Workers do not read top-level plan.md sections — /implement consumes per-phase **Knowledge context:** backlinks (Step 3.1 directive branch resolves them via resolve-manifest.sh into worker {{prior_knowledge}}). Distribution into per-phase Knowledge context happens in Step 5b #2 (concordance-assisted annotation) — see that step for the per-phase placement rule. Entries that don't bind to any specific phase still appear here; the manifest also catches them during review even when they have no per-phase home.
Keep this block at the full permissive surfaced set regardless of what binds to a task. The manifest is permissive; the weave into task lines is strict (Step 5b "Deliverable contract gate" — only scope-overlapping judgment-class norms become constraint clauses). A backlink staying here while its norm is also woven into a task is correct: the manifest is provenance, the task line is delivery.
The block is also a parse target: at close, the conformance renderer reads it as the spec-time discovery panel of closure-conformance.md and cross-tabulates each label against woven norms, recorded dispositions, and the shipped diff. Keep every bullet in the [[knowledge:...]] — annotation shape with a substantive annotation — a thinned or malformed manifest doesn't just weaken review, it blinds the closure read to norms nobody dispositioned.
Advisor declarations: For each matched skill whose domain overlaps with a phase's scope, consider adding an **Advisors:** entry. Set mode based on phase complexity — must-consult if the skill defines invariants workers must respect, on-demand otherwise.
Apply this contract after every terminal evaluator attempt in Steps 5a and 5.5. The evaluator supplies evidence; the lead decides the normalized protocol outcome. Never parse evaluator prose into a disposition.
Choose exactly one outcome: completed | failed | skipped | needs-decision. Preserve the evaluator's raw verdict byte-for-byte as --verdict. skipped and needs-decision require a reason; completed and failed forbid one. A registered evaluator that cannot execute is a filed skipped attempt, never a silent omission.
Build a version-1 evidence manifest with exactly these fields:
{
"schema_version": 1,
"evaluator_locator": "<skill or agent locator>",
"evaluator_template_version": "<12-char hash>",
"framework": "<active framework>",
"model": "<effective evaluator model>",
"final_round": 2,
"disposition_ledger_sha256": "<sha256 of the round/disposition ledger>",
"source_plan_sha256": "<sha256 of the plan the evaluator read>"
}
completed and failed require every evidence field. skipped and needs-decision keep every field present but may use explicit null when evidence is unavailable. Missing fields are errors, never defaults.
File the already-made judgment:
lore spec outcome "$SLUG" \
--ceremony <spec-design|spec-post-plan> \
--advisor "$EVALUATOR" --attempt-id "$ATTEMPT_ID" \
--outcome "$NORMALIZED_OUTCOME" --verdict "$RAW_VERDICT" \
--evidence-manifest "$EVIDENCE_JSON" [--reason "$REASON"]
Exact replay is idempotent; reusing an attempt id for different semantics is a refused collision. needs-decision may return status=partial when its auxiliary resolution row fails to append; exact retry recovers only that sink.
Ceremonies always run. No flag skips this step; no flag is required to run it. Don't ask whether to invoke them — invoke. Judgment applies to acting on the output, not to running the step.
EVALUATORS=$(lore ceremony get spec-design --work-item <slug>)
If non-empty JSON array, for each skill name in the array:
/<skill-name> <slug>
This evaluates the abstract plan. If WEAK or MISSING areas are identified, revise the abstract plan before proceeding to Step 5b. No evaluators are registered by default — opt-in via lore ceremony add spec-design <skill>.
After each evaluator reaches a terminal attempt, make the outcome judgment and file it under --ceremony spec-design using the ceremony outcome filing contract. A revision round receives a new attempt id; never overwrite the evidence identity of an earlier round.
When every evaluator holds a terminal disposition and any accepted revisions are persisted — including the no-evaluator case — a hosted session journals the design milestone:
if [[ -n "${LORE_SESSION_INSTANCE:-}" && -n "${LORE_SESSION_SLUG:-}" && -n "${LORE_SESSION_TYPE:-}" ]]; then
bash ~/.lore/scripts/session-step.sh \
--step-id spec:design --step-label "Design accepted" \
|| echo "[spec] Warning: design step not journaled; the persisted plan remains authoritative." >&2
fi
One row marks the accepted design state; evaluator attempts and individual revision rounds do not emit.
Draft concrete implementation sections on top of the approved abstract plan:
Intent anchor — if the work item has an intent_anchor in _meta.json, render a ## Intent Anchor section in plan.md immediately after ## Narrative and before ## Strategy/## Context. Write the anchor body verbatim from _meta.json.intent_anchor — no quoting, prefix label, or paraphrase. Before decomposing, name the tempting narrower implementation that would appear successful while violating the anchor; ensure the Goal, task constraints, and Verification cover the load-bearing promise or explicitly label the scope delta.
Follow the anchor with a **Scope delta:** line (default none — anchor preserved unchanged; if the spec narrows the capability, name the narrowing here) and a **Tempting narrower implementation:** heading the spec author fills in. The anchor body and **Scope delta:** line are verifier-enforced — the Step 5.6 gate refuses to regenerate tasks if missing or divergent. The **Tempting narrower implementation:** body is template-prescribed but not verifier-enforced — its presence forces the author to confront the failure mode, but the content is free-text that no parser can adjudicate. For work items without an intent_anchor field, omit the section entirely — the Step 5.6 verifier skips with a one-line stderr info message.
Phases — concrete implementation phases with tasks, file paths, objectives. Each phase includes **Knowledge context:**, **Tasks:** (checkbox lines), and optional **Retrieval directive:** / **Advisors:** / **Verification:** / **Split rationale:** (required when the phase has more than one task) / **Scope:** blocks.
Plan-as-unit rule. A plan is one phase by default. Each additional phase is a separate /implement worker batch with its own dispatch ceremony — write one phase per plan unless the split earns its keep across the entire run.
Split a plan into multiple phases only when all three conditions hold:
/implement dispatches concurrently. If file overlap forces every later phase sequential, generate-tasks.py chains them onto one worker — merging yields the same execution shape with less ceremony.If any condition fails, merge. The Phase-as-unit rule below still applies within the merged phase: most consolidated plans land at one phase, one task.
Phase-as-unit rule. A phase is the default delegation unit. Write one - [ ] checkbox per phase by default — deliver the phase objective across the listed files while honoring the phase design decisions. Each task spawns a fresh worker that loads its fixed context plus the phase brief; the four conditions below, not a flat per-task overhead, decide whether a split earns that spawn.
Split a phase into multiple tasks only when all four conditions hold:
generate-tasks.py onto one worker — the split buys nothing.If any condition fails, keep one task. Cross-phase dependencies are file-based. Uniform same-mechanism edits across many files stay one task — one worker doing one read-modify-write pass beats N fresh agents repeating the same edit.
Judgment class and the split calculus. Each task line carries a judgment class (mechanical | standard | judgment-dense; see the Deliverable contract gate below) that /implement routes to a worker-tier binding — mechanical to a cheaper model, judgment-dense to a stronger one. This gives the four conditions a second thing a split can buy beyond parallelism: separating a judgment-dense core from a mechanical shell over disjoint files, so each routes to its own tier instead of paying the strong-model rate across the whole phase. A judgment-density transition across disjoint files is a legitimate split point. But the four conditions still gate every split: a class mix that fails disjoint-file-ownership or real-parallel-execution stays one task, and a uniform same-mechanism sweep stays one task regardless of worker tier — a mechanical sweep never fragments into per-file tasks to shave model spend, because spawn overhead dwarfs the saving. The tipping point remains judgment density, not file count.
Context-envelope ceiling. A task's owned file set plus its phase brief must fit one worker's context envelope with working room to read, reason, and edit. When they do not, the split is forced — regardless of judgment class or whether the four conditions hold — because an over-large task passes every structural condition above yet fails in flight on worker-context exhaustion, the most expensive discovery point there is. This is a level-1 correctness-of-execution constraint (execution capacity), the sibling of the capability ceiling that keeps judgment-dense work off a model that cannot hold it — not a weighable fifth condition and not an aesthetic size threshold. It bounds task size from above exactly as no residue bounds it from below; between that floor and this ceiling, size stays structure-gated and empirically tuned via the (class, model, size) rework attribution — never trimmed to hit a size target.
Split rationale (required for multi-task phases). Any phase with more than one task carries a **Split rationale:** block — one or two sentences naming the judgment-density transition or genuine parallelism that earned the split. The Step 5.6 finalize gate refuses a multi-task phase that lacks it. A single-task phase omits the block.
Deliverable contract gate. Every task line names a durable artifact outcome — what gets built, refactored, authored, migrated, wired, or added. The valid primary verbs are: Implement / Refactor / Author / Migrate / Add support for / Wire... The following primary verbs are never tasks — they belong elsewhere: Verify / Check / Inspect / Run / Capture / Append / Cross-link / Note / Document-only.
A valid task line states the deliverable, the owned file or surface, and at least one design or integration constraint that scopes the worker's choices. Write the constraint's why in the codebase's own vocabulary — the actual symbol, flag, or error string it protects, not a paraphrase. Rationale is what a handoff loses first, and a worker who knows the reason can adapt when the letter of the constraint doesn't fit what the code turns out to be. When the task responds to a concrete failure, include the reproduction handle — the exact failing command, test case, or input — ranked above architectural narrative: a pointer the worker can run outranks a description of what it would show.
Judgment-class marker (required). Every task line ends with a trailing [class: mechanical | standard | judgment-dense] marker, placed after any [[knowledge:...]] backlinks so the deliverable verb stays first. It declares the judgment class — the worker tier /implement routes the task to:
worker.The class is explicit on every line — the Step 5.6 finalize gate refuses an unannotated task line. An unclassed line is not treated as standard: a legacy plan with no markers still regenerates (routing as plain worker), but re-finalizing it demands annotation.
Weave binding judgment-class norms into the constraint clause. When a preference/convention from Step 2 discovery binds to a task, render it as an imperative constraint clause in the task line itself — the instruction the worker executes — not only as a **Knowledge context:** backlink the worker must choose to fetch. The clause names the norm by its stable label (the entry slug/title the backlink resolves to) so the worker's Convention handling: report and the lead's completeness comparison reference the same identifier. Keep the backlink for provenance even when the norm is woven.
A norm binds only when both hold (strict weave):
related_files, file-path globs, ceremony scope, or activity domain intersect this task's owned files, objective, or surface.route-conventions-by-enforcement-class-delivery-vs. Judgment-class but no scope-overlap → backlink only (no constraint clause); neither → top-level manifest only.Weave only this binding subset — never the full permissive surfaced set, never mechanical norms. The top-level **Related preferences/conventions:** audit manifest stays unchanged (Step 5 — full permissive set); strict weaving into task lines is what prevents task-description bloat and dilution.
Example bound task line:
- [ ] Implement the retry wrapper in `src/net/client.py` — surface partial failures rather than swallowing them; honor `error-messages-name-the-failed-operation-and-the-fix` (name the failed operation and the corrective action in every raised error). [[knowledge:conventions/error-messages-name-the-failed-operation-and-the-fix]] [class: standard]
Route invalid units:
**Verification:** objective; do not duplicate into each task description.lore capture or notes.md); not worker work.Task format (intent+constraints). Default. State what the change accomplishes, what not to do, and what success looks like at the deliverable level. Opt into prescriptive format with **Task format:** prescriptive for mechanical work where step-by-step instructions are required.
Premise-wrong exit (standing, every phase). A phase brief hands its worker two sanctioned outcomes, not one: the deliverable, or the report "this cannot be built as scoped — here is what blocked me," naming the premise that failed contact with the code. The second is a first-class result from a colleague closer to the ground than the plan was; it routes to Step 6 follow-up investigation instead of forcing an approximation of a wrong plan. Never write a phase whose only expressible outcome is success.
Concordance-assisted annotation — after drafting phases, widen each phase's **Knowledge context:** block:
lore prefetch "<phase objective> <key file paths>" --type knowledge --limit 5 --scale-set=<bucket>
Declare --scale-set explicitly for every prefetch call. Missing declaration is an error.
Scale rubric — declare --scale-set explicitly on every prefetch. The four tiers (abstract, architecture, subsystem, implementation), boundary tests, multi-label encoding rules, and the ±1 query pattern live in skills/memory/SKILL.md Scale-Aware Navigation — read that section before declaring if the right bucket is not obvious. For decision-tree details see the classifier agent template (lore repo agents/classifier.md).
Add relevant entries as [[knowledge:...]] backlinks with "— why relevant" annotations. Investigation findings are the primary source; concordance is a widener.
Distribute surfaced preferences/conventions into per-phase Knowledge context (mandatory). The top-level **Related preferences/conventions:** block from Step 2 discovery is an audit manifest — it does not reach workers. To wire into worker {{prior_knowledge}}, distribute each surfaced entry into **Knowledge context:** of every phase whose scope plausibly overlaps:
related_files, file-path globs, ceremony scope, or activity domain intersects the phase's **Files:**, objective, or owned subsystem. Apply permissively — the Step 2 surfacing gate carries through to distribution. If a reviewer would expect the worker aware of the entry while editing the phase's files, distribute.[[knowledge:preferences/<entry>]] (or conventions/, or cross-cutting-conventions/) in **Knowledge context:**, with a worker-facing "— what to honor at implement time" annotation. Implementation-facing means tell the worker what to do, not just what it says./implement worker batch needing its own seeds. Do not consolidate across phases.**Related preferences/conventions:** — the manifest preserves the audit trail.**Knowledge context:**) carries every scope-overlapping entry, judgment-class or not, as a backlink — the worker can dismiss inapplicable ones. Weaving into a task's constraint clause (Deliverable contract gate above) is the strict subset: only the judgment-class entries that also scope-overlap the task become imperative constraint clauses naming the norm by its stable label. A judgment-class entry that binds to a task gets both — the backlink here (provenance + {{prior_knowledge}} seed) and the woven clause in the task line (the instruction the worker executes). A mechanical entry gets only the backlink (it is never woven)./implement Step 3.1 (directive branch) resolves seeds via resolve-manifest.sh from **Knowledge context:** backlinks + **Files:** paths. Distributing here flows entries through seeds → directive → worker {{prior_knowledge}} without new protocol surface in /implement.Retrieval directive derivation — after concordance widening, populate **Retrieval directive:** for each phase. Derivable from phase content alone; no user input.
Per-topic decomposition (v2 — default): the directive is a list of (topic, scale_set, [activity_vocab]) — exactly one focal topic plus up to five adjacent topics. Each topic fires its own BM25 OR query at its own scale_set; the worker prompt's ## Prior Knowledge block ends up sectioned (### Focal: <topic> / ### Adjacent: <topic>).
scale_set: subsystem,implementation. Seeds: phase's owned files (from **Files:**) plus [[knowledge:...]] entries in **Knowledge context:** about that subsystem. Prefer title-vocabulary terms (the entry's title tokens) over the topic label — title vocabulary resolves to entries the index can rank, while raw knowledge:... strings tokenize as a single literal and miss the index.scale_set one tier above focal's bottom — typically architecture,subsystem. Seeds: title-vocabulary terms from canonical entries about that adjacent subsystem — not the topic label, not the focal seeds. Weak seeds (right scale, wrong entries) are the dominant failure mode — re-derive from adjacent entries' titles and resolved paths.$KDIR/_meta/activity-vocab.yaml by matching its file-path globs against the topic's owned files; do not invent activity tokens inline. The activity-vocab file is the single authority. When present, the topic fires one extra BM25 OR query at the same scale_set with these tokens (query_kind=activity).role: focal entry. Zero-focal or multi-focal v2 is a hard parse error in generate-tasks.py — not silently accepted, not normalized to legacy. If no genuine focal candidate emerges (e.g., purely cross-cutting refactor), emit the legacy flat directive — that path remains valid for rollout compatibility.Seeds derivation (mandatory): per topic, collect from two sources — (a) [[knowledge:...]] backlinks in **Knowledge context:** whose subject matches the topic (resolve to entry title and path-vocabulary terms before emitting — raw knowledge: strings won't tokenize); (b) **Files:** paths the topic owns (verbatim). Deduplicate per topic. Empty seed union → the topic itself is suspect; drop it rather than emit an empty seeds: bullet.
Defaults: hop_budget: 1. Per-section limits: focal limit: 8, adjacent limit: 4 (tunable). scale_set: is mandatory per topic; omitting is an error. Pick abstract, architecture, subsystem, or implementation; multi-label form (e.g., architecture,subsystem) is allowed for adjacent pairs. Omit filters: unless type or category filtering adds value.
Format (v2 — default):
retrieval_directive:
version: 2
topics:
- role: focal
topic: "<short label>"
seeds:
- "[[knowledge:path#heading]]"
- "path/to/owned/file.py"
scale_set: [subsystem, implementation]
activity_vocab: [pytest, fixture, assertion, mock] # optional; from _meta/activity-vocab.yaml
limit: 8
- role: adjacent
topic: "<adjacent subsystem label>"
seeds:
- "<title-vocabulary terms from a canonical entry about the adjacent subsystem>"
scale_set: [architecture, subsystem]
limit: 4
# ...up to 5 adjacent total
hop_budget: 1
Format (legacy flat — rollout compatibility):
**Retrieval directive:**
- seeds: [[knowledge:path#heading]], path/to/file.py, ...
- hop_budget: 1
- scale_set: <bucket>
The legacy form continues to resolve to a single focal topic at the declared scale_set so existing plans don't break.
Omission rule: if a phase has neither **Knowledge context:** backlinks nor **Files:** entries, omit the **Retrieval directive:** block and add a comment: <!-- no directive: no backlinks or files to derive seeds from -->.
Position: place **Retrieval directive:** immediately after **Knowledge delivery:** (or after **Files:** / **Objective:** when **Knowledge delivery:** is absent) and before **Knowledge context:**.
Open Questions — anything investigations couldn't resolve.
Present the synthesized plan to the user for review.
lore work regen-tasks <slug>
Inspect phase_cost_summary as a sanity check — a single task far larger than its peers may signal an under-decomposed deliverable worth a closer read. Cost diagnostics are advisory only; the Plan-as-unit rule, Phase-as-unit rule, and Deliverable contract gate in Step 5b are the binding gates. Do not split tasks merely because they fall above an avg-comparison threshold, and do not merge tasks merely because they fall below one. The avg-comparison heuristic is post-hoc and uniform-thinness blind; trust the intrinsic gates instead.
bash ~/.lore/scripts/verify-plan-backlinks.sh "$WORK_DIR/<slug>/plan.md" "$KNOWLEDGE_DIR" --fix
Output: {verified: N, corrected: [...], unresolved: [...]}.
[broken backlink] bullets.The Step 5.6 finalize verb re-runs backlink verification terminally; this early pass exists to surface broken links before the Step 5.1 review, not to replace the terminal check.
For each phase, run lore search "<phase objective keywords>" --scale-set subsystem,implementation --limit 3. If results exist but the phase has no **Knowledge context:** block, add the most relevant entry as a backlink with an implementation-facing annotation.
Before finalizing, present 5-10 bullet points covering key assumptions, behavioral claims (mark [verified] or [unverified]), design decisions with rejected alternatives, scope boundaries, and any unresolved backlinks.
Format:
Before finalizing this plan, here is my understanding of the key assumptions:
- [verified] <claim> → Investigation: <topic>, Assertion #N
- [unverified] <claim> → Investigation: <topic>, Assertion #N
- <decision statement> (over <rejected alternative>) → Design Decision: D1: <title>
- <scope boundary> → Goal / user input
- [intent anchor] <anchor body verbatim from `_meta.json.intent_anchor`> — **Scope delta:** <none — anchor preserved unchanged | named narrowing> → Step 5b item 0 intent-anchor preservation (omit this bullet when the work item has no `intent_anchor`)
- [broken backlink] [[knowledge:path]] could not be resolved → Step 5.0a backlink check
...
Does this match your understanding? Any corrections?
Gate: Do not proceed to Step 5.3 until the user explicitly confirms or provides corrections. If --yes, skip (auto-proceed).
→ trace links.plan.md.Updated understanding after your corrections:
- [corrected] <revised claim> → <source>
...
Anything else to adjust?
Before finalizing, present the plan phases as structured summaries. This is a separate gate from Step 5.1 — that validates understanding; this validates the work plan.
Phase N: <Name>
Objective: <what this phase accomplishes>
Mechanism: <HOW — specific technical approach, 1-3 sentences>
Scope: <files and components touched>
Tasks: <N tasks>
Workers: N (max concurrent from task DAG topology)
Read recommended_workers from tasks.json. End with: Review the phases above. Approve to proceed, or request changes.--yes, skip (auto-approve).plan.md, re-present only changed summaries. Repeat until approved./spec <slug>.Invoke /remember scoped to the spec investigation. Always invoke it — even when no observation appears to meet the gate. The gate lives in /remember; rejecting candidates is /remember's job, not the lead's. Pre-filtering observations because "nothing qualifies, so /remember would be a no-op" is the bypass shape named in the commitment protocol. A run that captures zero entries is a valid terminal so long as /remember actually evaluated the observations.
Capture posture: generous in, exacting in form. Verification happens downstream, at consumption — every entry meets real code when a later agent reads it, and entries that fail that contact get pruned. So the expensive mistake is a malformed entry, not an extra one: don't spend effort predicting whether a future session will need an insight; spend it putting the insight in the form that lets that session retrieve and check it. Every capture written or promoted here takes the five-part form:
file:line@SHA anchoring the claim; this is the anchor lore verify checks.Every lore capture call must carry provenance flags; for captures promoted from researcher observations, preserve the original producer's attribution:
--producer-role spec-lead --protocol-slot Synthesis --work-item <slug> --template-version $LEAD_TEMPLATE_VERSION--producer-role researcher --capturer-role spec-lead --source-artifact-ids <researcher-report-ids> --protocol-slot Synthesis --work-item <slug> --template-version $RESEARCHER_TEMPLATE_VERSION/remember Research findings from <work item title> — Read all **Observations:** entries from investigation reports in plan.md and evaluate each: mechanism-level patterns, design rationale, and structural footprint signals all qualify; implementation facts already expressed in Tier-2 assertions do not. Also capture cross-investigation synthesis patterns not surfaced individually. Shape every capture in the five-part form: contrastive pair, the codebase's own vocabulary, explicit applicability trigger, surprise-derived content, provenance pointer.
Apply the provenance flags above on every `lore capture`.
Ceremonies always run before terminal finalization. No flag skips this step; no flag is required to run it. Judgment applies to acting on output, not to whether the registered obligation executes.
EVALUATORS=$(lore ceremony get spec-post-plan --work-item <slug>)
Invoke every registered evaluator. Present its output to the user. If the lead accepts changes, revise plan.md, repeat the affected review gates, and run a fresh evaluator attempt. After each terminal attempt, make the normalized outcome judgment and file it under --ceremony spec-post-plan using the ceremony outcome filing contract.
Do not finalize while a post-plan result still requires a plan edit or human decision. needs-decision is durable evidence of that open judgment, not permission to route around it.
Run Step 5.6's three lead-owned preflight asserts now, without finalizing: sweep the plan's named paths, symbols, and packages against the live tree (marking unresolvable names (unverified)), check every instructed invocation against the live script, and check that Tier-2 emission instructions point to the canonical validator contract. If any assert — or the finalize verb itself — refuses, fix plan.md and re-run the affected ceremony before Step 5.6 invokes lore spec finalize.
With the post-plan ceremony terminal and all three preflight asserts passing, a hosted session journals the plan-ready milestone before entering Step 5.6:
if [[ -n "${LORE_SESSION_INSTANCE:-}" && -n "${LORE_SESSION_SLUG:-}" && -n "${LORE_SESSION_TYPE:-}" ]]; then
bash ~/.lore/scripts/session-step.sh \
--step-id spec:plan-ready --step-label "Plan ready" \
|| echo "[spec] Warning: plan-ready step not journaled; the persisted plan and preflight result remain authoritative." >&2
fi
Finalization emits no step of its own — lore spec finalize keeps the later, distinct terminus_reached row, and a refused preflight or finalize synthesizes no step history.
Lead-owned preflight (three prose asserts the verb cannot run): validate the plan against live sources, never from memory: (1) every path, symbol, and package the plan names resolves against the live tree — a mechanical grep or lookup per name; mark anything that doesn't resolve (unverified) where it appears, so the marking itself is the falsifier a later reader checks; (2) any script invocation block the plan instructs agents to run must match the live script's current flags (check --help or source — script schemas drift faster than plans); (3) Tier-2 emission instructions point at the validator's canonical required-field set rather than enumerating fields inline. Fix plan.md first if any fails — a deterministic script can assert JSON structure, but adjudicating prose against live sources is the lead's judgment.
Then close the plan through the finalize verb. Do not hand-run its composed checks or writers; finalize owns backlink verification, the intent-anchor hard gate, task regeneration, healing, retrieval-directive assertions, its spec-verb atom, and last-write telemetry:
lore spec finalize <slug>
Show the verb's output. The anchor gate enforces structural anchor preservation and scope-delta attestation, not semantic non-drift — semantic alignment between the anchor and the rest of the plan remains a spec-author responsibility, with downstream reviewers (e.g., /codex-plan-review) as the semantic backstop. A no-anchor work item reports the gate as skipped with the verifier's reason (absence is legible, not silent).
Refusal handling:
**Scope delta:** missing; contract failures name the failing phase), fix plan.md, and re-run lore spec finalize <slug> until it passes.A refused finalize emits no telemetry row and no spec-verb atom; re-running after a fix appends a fresh point-event row per run — expected, not duplication.
If gaps are identified (from evaluator feedback or user review):
lore spec open again, and execute the returned directives.Run lore work heal after any changes.
After finalization, suggest:
Consider `/retro <slug>` to evaluate knowledge system effectiveness for this spec.
When emitting plan.md in Step 5b (including item 0's Intent Anchor render), read skills/spec/templates/plan.md for the canonical plan structure. The sidecar holds the full fenced template — Goal, Narrative, Intent Anchor, Strategy, Context, Investigations, Design Decisions, Architecture Diagram, Phases, Open Questions, Related — with the inline HTML-comment guidance preserved alongside each section it governs.