- name
- create-walkthrough
- description
- Collaborative argumentative walkthrough for complex implementations. REQUIRES /interview (user context) and /ask consult (persona review) BEFORE writing. Combines claim verification, Mermaid diagrams, structured tables, and adversarial human review into a prosecution brief.
- allowed-tools
- Bash, Read, Write, WebFetch
- triggers
- ["create walkthrough","write walkthrough","walkthrough","honest walkthrough","implementation walkthrough","why will this work","explain the implementation","walk me through","pre-launch review"]
- metadata
- {"short-description":"Collaborative walkthrough with claim verification"}
- provides
- ["create-walkthrough"]
- composes
- ["task-monitor","agentic-evals"]
- disciplines
- ["human-collaboration","content-creation"]
# create-walkthrough
Generate honest, argumentative walkthrough documents for complex implementations.
Not a status report or handoff document. A **prosecution brief** where the agent
argues why an implementation should succeed, admits what could go wrong, and the
user pokes holes.
**This is a COLLABORATIVE skill.** The agent does NOT write a walkthrough alone.
## Why This Exists
A walkthrough caught two real bugs before a pipeline launch:
1. A false claim about a missing dependency (agent wrote it, agent believed it, user caught it)
2. A missing semantic quality check the deterministic assessment couldn't provide
The value isn't the document structure. It's:
- **Collaboration**: the human and a persona expert contribute BEFORE writing starts
- **Risk-forcing**: every change MUST have "what could still go wrong"
- **Claim verification**: every factual statement is audited against actual code
- **User review surface**: the document exists so the human can push back
### Why Collaboration Is Non-Negotiable
An agent writing a walkthrough alone produces a **monologue** — it explains its own
work to itself. The agent's blind spots become the walkthrough's blind spots. Real
bugs were caught in the episodic-archiver v2 walkthrough not by the agent, but by the
user reading critically. The interview and persona consultation exist to surface
concerns the agent **cannot see**.
**Incident (2026-02-13):** Agent skipped interview + persona consultation for the
episodic-archiver v2 walkthrough. Result: a technically correct but one-dimensional
document that missed the user's concern about conversation prediction classifiers and
the persona's expertise in user behavioral modeling. The walkthrough failed at its
primary purpose — being a collaboration surface.
## When to Use
Use `/create-walkthrough` when ALL of these are true:
- The system has **failed before** (at least one prior attempt)
- The implementation is **complex** (multi-file, multi-concern)
- You're about to **launch or deploy** (not still designing)
- The user needs to **review and approve** before proceeding
Do NOT use for:
- First-time implementations (use `/plan` instead)
- Simple features or bug fixes
- Agent-to-agent handoff (use `/create-context` instead)
- General project health (use `/assess` instead)
## How It Differs
| Skill | Modality | Question Answered |
|-------|----------|-------------------|
| `/create-context` | Descriptive | "What happened? What's the state?" |
| `/assess` | Evaluative | "Is this healthy? What's broken?" |
| `/plan` | Prescriptive | "What should we do next?" |
| **`/create-walkthrough`** | **Argumentative + Collaborative** | **"Why should this work when previous attempts failed?"** |
---
## Pre-Flight Checklist (BLOCKING)
Before writing ANY walkthrough content, verify ALL of these:
| Gate | Requirement | How to Complete |
|------|-------------|-----------------|
| **Interview** | User has answered questions about failures, concerns, scope | Use `/interview` or AskUserQuestion |
| **Persona** | A domain persona has reviewed the changes | Use `/ask consult <persona>` |
| **Memory** | Prior failures/lessons recalled | Use `/memory recall` |
| **Code read** | Agent has read the actual implementation files | Use Read tool |
**If ANY gate is incomplete, STOP. Do not write the walkthrough.**
The agent MUST NOT rationalize skipping gates:
- "I have deep session context" is NOT a reason to skip the interview
- "No persona is relevant" is NOT true — every implementation has a domain expert
- "The user didn't ask for persona input" is NOT relevant — the skill requires it
---
## Workflow
### Phase 1: Human Interview (MANDATORY — NO EXCEPTIONS)
**ALWAYS ask the human.** Even if you implemented the code yourself in this session.
Even if you think you know the answers. The human sees things you don't.
Use `/interview` or `AskUserQuestion` to gather:
```json
[
{
"id": "failures",
"text": "What has failed in previous attempts? List specific failure modes.",
"type": "text",
"header": "Failures"
},
{
"id": "concerns",
"text": "What are you most worried about this time?",
"type": "text",
"header": "Concerns"
},
{
"id": "constraints",
"text": "What deployment constraints apply?",
"header": "Constraints",
"options": [
{"label": "Single process only", "description": "No concurrent daemons"},
{"label": "Must survive API outages", "description": "External dependency resilience"},
{"label": "Unattended overnight", "description": "No human monitoring"},
{"label": "Resource constrained", "description": "Memory/CPU/VRAM limits"}
],
"multi_select": true
},
{
"id": "scope",
"text": "Which files/systems should the walkthrough cover?",
"type": "text",
"header": "Scope"
},
{
"id": "persona",
"text": "Which persona should review this? (Pick the domain expert most relevant to this system.)",
"type": "text",
"header": "Reviewer"
}
]
```
**Why this can't be skipped:** The human's concerns shape the walkthrough's focus. Without
asking, the agent writes about what IT thinks matters. The episodic-archiver v2 walkthrough
missed the user's interest in conversation prediction classifiers because the agent never
asked. The interview is how the human steers the walkthrough.
**Minimum interview:** If `/interview` is unavailable, use `AskUserQuestion` with at
minimum these 3 questions:
1. "What are you most worried about with this implementation?"
2. "What should the walkthrough focus on — what do you need to be convinced of?"
3. "Which persona should review this? (e.g., Embry for user modeling, Brandon for SPARTA, Margaret for extraction)"
Also gather from automated sources:
- `/memory recall` for past failures, lessons, and assessments related to this system
- `git log` for recent changes and commit messages
- `CONTEXT.md` for current state documentation
### Phase 1b: Persona Consultation (MANDATORY — NO EXCEPTIONS)
**ALWAYS consult a persona.** The user nominates one in the interview (Phase 1). If the
user didn't specify, pick the most relevant domain expert yourself and confirm with the
user: "I'll consult [Persona] — they have expertise in [domain]. Sound right?"
Use `/ask consult <persona>` with a summary of changes:
```
We're about to deploy [system]. Here's what changed:
1. [Change 1 — one sentence]
2. [Change 2 — one sentence]
3. [Change N — one sentence]
What concerns you? What are you satisfied with? What would you watch for
in the first hour of deployment?
```
**Why this can't be skipped:** Different personas surface different concerns. The agent
may not realize that a design pattern is risky in a specific domain — but the persona
will. Examples:
| Persona | What They'd Catch That the Agent Wouldn't |
|---------|------------------------------------------|
| **Embry** | User behavioral modeling gaps, conversation prediction feasibility, linguistics edge cases |
| **Brandon Bailey** | SPARTA-specific: grounding formula gaps, framework term coverage, D3FEND abstraction levels |
| **Margaret Chen** | Extraction quality: PDF parsing failures, table detection false positives, data integrity |
| **Horus Lupercal** | System architecture: single points of failure, resilience under adversarial conditions |
**The persona's output becomes the "Expert Commentary" section of the walkthrough.**
```markdown
## Expert Commentary
**[Persona Name]** — [Role/Title]
> **What I'm satisfied with:**
> - [Specific thing persona approves, with domain reasoning]
> - [Another]
>
> **What concerns me:**
> - [Specific concern, grounded in persona's expertise]
> - [Another]
>
> **What I'd watch for in the first hour:**
> - [Observable metric or behavior the persona would monitor]
```
This transforms the walkthrough from "agent explains agent's work" to "domain expert
reviews agent's work." The persona brings knowledge the agent may lack.
**Rule:** The persona consultation is GENERIC. Any persona from `personas.yaml` can be
consulted. Do NOT build persona-specific logic into the skill.
### Phase 2: Analyze the Implementation
**Only proceed here after BOTH Phase 1 and Phase 1b are complete.**
Read the actual code. For each significant change:
1. **Identify what it replaces** (the old approach that failed)
2. **Understand the mechanism** (how the new code works, line numbers)
3. **Find the integration points** (where it connects to existing code)
4. **Assess the risk** (what could go wrong with this specific change)
5. **Cross-reference with interview** (does this address the user's concerns?)
6. **Cross-reference with persona** (does this address the persona's concerns?)
### Phase 3: Write the Walkthrough
Use this structure. **All sections are REQUIRED.**
The walkthrough MUST incorporate:
- User's concerns from the interview (Phase 1)
- Persona's concerns and satisfactions from the consultation (Phase 1b)
- Memory recall results showing prior failures and lessons
```markdown
# [System Name] v[N]: Honest Walkthrough
**Date:** YYYY-MM-DD
**File(s):** `path/to/main/file.py` (N lines)
**Status:** [Preflighted / Tested / Production-tested]
**Reviewed by:** [Persona Name] ([Role])
**User concerns addressed:** [List from interview]
---
## Why Previous Versions Failed
### Failure 1: [Short Title]
**What we did:** [Factual description of the approach]
**Why it failed:** [Root cause, not symptoms]
### Failure N: ...
---
## What v[N] Changes
### Change 1: [Short Title] (lines X-Y)
[Description of the change with code snippets]
**What this fixes:** [Which failure mode from above]
**What could still go wrong:** [Honest risk — REQUIRED, cannot be empty]
**Honest risk level:** LOW / MEDIUM / HIGH — [justification]
### Change N: ...
---
## Expert Commentary
**[Persona Name]** — [Role/Title]
> **What I'm satisfied with:**
> - [From Phase 1b consultation]
>
> **What concerns me:**
> - [From Phase 1b consultation]
>
> **What I'd watch for in the first hour:**
> - [From Phase 1b consultation]
---
## Data Flow Diagram
[Use /create-figure with Mermaid backend to generate a flowchart]
```mermaid
flowchart TD
A[Step 1] --> B[Step 2]
B --> C{Decision}
C -->|Yes| D[Path A]
C -->|No| E[Path B]
```
---
## Risk Matrix
[Use markdown table — /create-table if PDF output needed]
| Change | Fixes | Risk | Observable Failure |
|--------|-------|------|--------------------|
| ... | ... | LOW/MED/HIGH | How you'd know it broke |
---
## Remaining Risks (Honest Assessment)
### Risk 1: [Title] (SEVERITY)
[Description, mitigation, what would actually fix it]
---
## What Success Looks Like
| Metric | Healthy | Warning | Sick |
|--------|---------|---------|------|
| ... | ... | ... | ... |
---
## How to Launch / Monitor / Kill
[Exact commands — copy-pasteable]
---
## Bottom Line
**Will it work?** [Honest one-paragraph assessment]
**What's genuinely different this time?** [Numbered list]
**What's the same?** [What DIDN'T change — often reveals the real bottleneck]
---
## Next Steps — Your Call
[If there are open questions or branching next steps, include interview-style
questions so the user can steer what happens next. Use numbered options with
descriptions. These should be REAL decisions, not rubber-stamp confirmations.]
**1. [Decision question]**
- a) [Option] — [what this means, tradeoff]
- b) [Option] — [what this means, tradeoff]
- c) [Option] — [what this means, tradeoff]
**2. [Another decision]**
- a) ...
- b) ...
[For HTML walkthroughs, render these as interactive elements if possible.
For markdown, use the numbered format above so the user can reply "1b, 2a".]
```
### Phase 4: Claim Verification (CRITICAL)
Before presenting the walkthrough to the user, run the claim verification engine:
```bash
./run.sh verify --file path/to/walkthrough.md
```
The verifier extracts and checks:
| Claim Type | Example | Verification |
|-----------|---------|-------------|
| **File paths** | "`src/foo.py` (3,337 lines)" | File exists, line count matches |
| **Function names** | "`assess_qra()` on line 275" | Function exists at that line |
| **Package availability** | "`sentence_transformers` not installed" | Check pyproject.toml, pip list, venv |
| **Environment vars** | "`EMBEDDING_PORT` defaults to 8602" | Grep code for the default |
GitHub에서 보기