- name
- e2e-template-testing
- description
- End-to-end validation of coordinator and agent template changes
- domain
- development
- confidence
- high
- source
- manual
## Context
Squad's coordinator prompt (`squad.agent.md`) and agent charters (e.g.
`scribe-charter.md`) are shipped as templates in `.squad-templates/`. Changes to
these files affect how every squad session behaves — but unit tests can't catch
prompt-level regressions because the prompts are interpreted by an LLM at
runtime.
This skill describes how to validate template changes end-to-end by running real
squad sessions against a locally-built CLI that includes your modified templates.
## When To Use
- You changed `.squad-templates/squad.agent.md` (coordinator prompt)
- You changed `.squad-templates/scribe-charter.md` or other agent charters
- You changed `.squad-templates/notes-protocol.md` or helper scripts
- You added new conditional blocks (e.g. state-backend-aware spawn templates)
- You modified the init scaffolding that writes templates to target repos
## Prerequisites
- **Node.js** ≥20, **npm** ≥10
- **Git** CLI
- **GitHub Copilot CLI** (`copilot` or `ghcs`) installed
- A local clone of the squad repo on your feature branch
## Workflow
### Step 0 — Post initial tracking comment (**FIRST action — before anything else**)
If `PR_NUMBER` and `REPO` are both set, **the absolute first thing you do** — before
fast-fail checks, before building, before creating any repos — is post the initial
tracking comment with all steps marked as `:hourglass_flowing_sand: Pending`.
This gives reviewers immediate visibility that a run is in progress and what to expect.
```powershell
$runStart = Get-Date
$body = @"
## E2E Progress - PR $env:PR_NUMBER
| Step | Status | Started | Duration |
|---|---|---|---|
| 1. Fast-fail checks (build :cd: link :cd: ``squad version``) | :hourglass_flowing_sand: Pending | --:-- | -- |
| 2. Create test repo(s) | :hourglass_flowing_sand: Pending | --:-- | -- |
| 3. ``squad init`` + file verification | :hourglass_flowing_sand: Pending | --:-- | -- |
| 4. Run sessions | :hourglass_flowing_sand: Pending | --:-- | -- |
| 5. Verify outcomes | :hourglass_flowing_sand: Pending | --:-- | -- |
| 6. Record verdicts + post final comment | :hourglass_flowing_sand: Pending | --:-- | -- |
| Symbol | Meaning |
|---|---|
| :hourglass_flowing_sand: | Not started |
| :arrows_counterclockwise: | Running |
| :white_check_mark: | Passed |
| :x: | Failed |
| :warning: | Passed with caveats |
*Run started: $($runStart.ToString('HH:mm')) — all steps pending*
"@
$tmpFile = [System.IO.Path]::GetTempFileName()
$utf8NoBom = New-Object System.Text.UTF8Encoding $false
[System.IO.File]::WriteAllText($tmpFile, $body, $utf8NoBom)
$response = gh api "repos/$env:REPO/issues/$env:PR_NUMBER/comments" --method POST --field "body=@$tmpFile" | ConvertFrom-Json
$env:COMMENT_ID = $response.id
Remove-Item $tmpFile -Force
Write-Host "Progress comment posted — ID: $($response.id)"
```
After posting, immediately update the comment to mark Step 1 as `:arrows_counterclockwise: Running` (do NOT wait — this is a two-step sequence: post all-pending, then immediately update to Step 1 running). Then proceed to Step 1.
See **Progress Reporting** for the full comment lifecycle and update patterns.
**⚠️ STOP before continuing:** If posting the comment fails (network error, auth error), abort the run and report the failure. Do not proceed silently without a tracking comment.
### Step 1 — Build the CLI from your branch
```bash
cd /path/to/squad # your feature branch
npm install
npm run build -w packages/squad-sdk && npm run build -w packages/squad-cli
# Link so `squad` command uses your local build (workspace flag — no cd required)
npm link -w packages/squad-cli
```
Verify: `squad version` output includes the `-preview` suffix (e.g., `x.y.z-preview`),
confirming the local dev build is active. If the output shows a plain semver without
`-preview`, the globally-installed npm package is still in use — re-check the link step.
See [CONTRIBUTING.md — Making the `squad` Command Use Your Local Build](../../../CONTRIBUTING.md)
for the full guidance on local dev versioning.
### Step 2 — Create a disposable test repo
```bash
mkdir /tmp/sq-test-1 && cd /tmp/sq-test-1
git init
echo "# Test Project" > README.md
echo '{"name":"test-project","version":"1.0.0"}' > package.json
mkdir src
echo "export function hello() { return 'world' }" > src/index.ts
git add -A && git commit -m "init: test project"
```
Keep the project small — you only need enough for the coordinator to recognize a
codebase and hire a team.
### Step 3 — Init a squad with your modified templates
```bash
squad init
# If testing a specific feature (e.g. state backends):
# squad init --state-backend git-notes
```
Verify the init produced the expected files:
```bash
ls -la .squad/
cat .squad/team.md # should have ## Members with 3+ agents
cat .squad/config.json # should reflect any CLI flags you passed
```
### Step 4 — Run a real session and capture output
Use the Copilot CLI's `-p` flag with `--allow-all-tools` for non-interactive sessions.
`--allow-all-tools` is **required** for automated/non-interactive runs — without it,
tool calls (including file writes) prompt for confirmation and block.
```powershell
# PowerShell (Windows)
copilot --agent squad --allow-all-tools -p "Picard, decide what testing framework to use. Write your decision." `
2>&1 | Tee-Object evidence/session-task.log
```
```bash
# Bash (macOS/Linux)
copilot --agent squad --allow-all-tools -p "Picard, decide what testing framework to use. Write your decision." \
2>&1 | tee evidence/session-task.log
```
Alternatively, set the `COPILOT_ALLOW_ALL=1` environment variable instead of the flag.
For multi-turn workflows, run sequential sessions:
```powershell
# Session A: give the team a task
copilot --agent squad --allow-all-tools -p "prompt A" 2>&1 | Tee-Object evidence/session-A.log
# Session B: verify state persisted
copilot --agent squad --allow-all-tools -p "What decisions has the team made?" 2>&1 | Tee-Object evidence/session-B.log
```
### Step 5 — Verify the outcome
Check that your template change had the expected effect. Common checks:
```bash
# State location (for state-backend changes)
git notes --ref=squad list # git-notes backend
git ls-tree -r squad-state # orphan backend
ls .squad/agents/*/history.md # worktree backend
# Coordinator behavior (grep session log)
grep "STATE_BACKEND" evidence/session-task.log
grep "spawn" evidence/session-task.log
# File tree diff
git diff --stat HEAD~1 # what changed on working branch
git log --all --oneline # commits across all branches
```
### Step 6 — Record the verdict
Create an `evidence/verdict.md` in each test repo:
```markdown
## Test: [scenario name]
**Backend:** worktree | git-notes | orphan | two-layer
**Branch:** [your feature branch]
**Result:** PASS | PARTIAL | FAIL
**Duration:** Xm Ys
### What was verified
- [ ] Coordinator identified feature correctly (from session log)
- [ ] Agent was spawned via `task` tool (not simulated)
- [ ] team.md has ## Members with 3+ agents
- [ ] State landed in correct location
- [ ] No unexpected side effects
### Evidence files
- session-task.log — full session output
- git-log.txt — `git log --all --oneline`
### Notes
[anything unusual or noteworthy]
```
Record the wall-clock time from the start of Step 1 (fast-fail checks) to the end
of Step 6 (verdict posted). This is the full E2E run duration for this scenario.
## Progress Reporting
Use this section only when you are running E2E validation for an open PR. If
`PR_NUMBER` and `REPO` are both set, post and maintain a live tracking comment
in the PR thread. If either value is missing (for example, a local-only run),
skip progress reporting silently.
### Start the tracking comment (Step 0 — see Workflow above)
The initial comment must be posted as **Step 0** — the absolute first action before
anything else. See the Step 0 block in the Workflow section for the exact code.
The subsections below describe how to **update** the comment at each step boundary.
For reference, here is the initial all-pending comment body posted in Step 0:
1. Post a PR comment before Step 1 begins:
```bash
gh pr comment "$PR_NUMBER" --repo "$REPO" --body "## E2E Progress\n\n| Step | Status | Started | Duration |
|---|---|---|---|
| 1. Fast-fail checks (build · link · \\`squad version\\`) | ⏳ Pending | --:-- | -- |
| 2. Create test repo(s) | ⏳ Pending | --:-- | -- |
| 3. \\`squad init\\` + file verification | ⏳ Pending | --:-- | -- |
| 4. Run sessions | ⏳ Pending | --:-- | -- |
| 5. Verify outcomes | ⏳ Pending | --:-- | -- |
| 6. Record verdicts + post final comment | ⏳ Pending | --:-- | -- |
\n| Symbol | Meaning |
|---|---|
| ⏳ | Not started |
| 🔄 | Running |
| ✅ | Passed |
| ❌ | Failed |
| ⚠️ | Passed with caveats |"
```
2. Capture the comment ID immediately after posting it:
```bash
COMMENT_ID=$(gh api "repos/$REPO/issues/$PR_NUMBER/comments" --jq '.[-1].id')
```
3. Treat Step 1 as in progress as soon as the comment exists. Update the body so
Step 1 shows `🔄 Running` and every later step remains `⏳ Pending`.
### Update the tracking comment after every step boundary
1. When marking a step `🔄 Running`, record `$startTime = Get-Date` and store the
`HH:MM` start time in that row's **Started** column.
2. Edit the existing comment in place; do not post a new progress comment:
```bash
gh api --method PATCH "repos/$REPO/issues/comments/$COMMENT_ID" --field body="..."
```
3. When marking a step `✅`, `❌`, or `⚠️`, compute
`$duration = (Get-Date) - $startTime` and format it as
`"{0}m {1}s" -f [int]$duration.TotalMinutes, $duration.Seconds`.
4. Update the completed step row to `✅`, `❌`, or `⚠️`, keep its original
`HH:MM` value in **Started**, and write the formatted duration in **Duration**.
5. Keep all previously completed rows unchanged.
6. Mark the next step as `🔄 Running` and set its **Started** value.
7. Leave later steps as `⏳ Pending` with `--:--` for **Started** and `--` for
**Duration**.
8. If a step fails and you stop early, still update the comment so the failed step
shows `❌` with its original start time and computed duration, and Step 6
becomes `🔄 Running` while you prepare the final verdict.
### Use this status legend in the comment
| Symbol | Meaning |
|---|---|
| ⏳ | Not started |
| 🔄 | Running |
| ✅ | Passed |
| ❌ | Failed |
| ⚠️ | Passed with caveats |
### Use exact step names and order
Keep these six rows in this exact order every time you update the comment:
1. Fast-fail checks (build · link · `squad version`)
2. Create test repo(s)
3. `squad init` + file verification
4. Run sessions
5. Verify outcomes
6. Record verdicts + post final comment
### Handle Windows comment bodies safely
On Windows PowerShell 5.1, use the **`--field body=@file`** pattern to post comment
bodies. Write the content to a temp file using UTF-8 **without BOM**, then pass
`--field "body=@$tmpFile"` to `gh api`. This is more reliable than piping JSON
through `--input -` on PS 5.1, which can silently corrupt multi-byte characters
even with `[Console]::OutputEncoding = UTF8`.
Key rules:
- Use `New-Object System.Text.UTF8Encoding $false` (the `$false` disables the BOM).
`[System.Text.Encoding]::UTF8` writes a BOM which GitHub renders as a stray
character (``) at the start of the comment.
- Use `--field "body=@$tmpFile"`, NOT `--input -` or `--input filename`, for
comment body updates. The `@` prefix tells `gh` to read the field value from
the file rather than treating the path as a literal string.
- Clean up the temp file after posting.
- Scrub any local absolute paths from the body before posting (see PII Protection
section).
```powershell
$step1StartTime = Get-Date
$step1Started = $step1StartTime.ToString('HH:mm')
$step1Duration = (Get-Date) - $step1StartTime
$step1DurationText = "{0}m {1}s" -f [int]$step1Duration.TotalMinutes, $step1Duration.Seconds
$step2StartTime = Get-Date
$step2Started = $step2StartTime.ToString('HH:mm')
$body = @"
## E2E Progress
| Step | Status | Started | Duration |
|---|---|---|---|
| 1. Fast-fail checks (build · link · `squad version`) | :white_check_mark: Passed | $step1Started | $step1DurationText |
| 2. Create test repo(s) | :arrows_counterclockwise: Running | $step2Started | -- |
| 3. `squad init` + file verification | :hourglass_flowing_sand: Pending | --:-- | -- |
| 4. Run sessions | :hourglass_flowing_sand: Pending | --:-- | -- |
| 5. Verify outcomes | :hourglass_flowing_sand: Pending | --:-- | -- |
| 6. Record verdicts + post final comment | :hourglass_flowing_sand: Pending | --:-- | -- |
| Symbol | Meaning |
Ver en GitHub