Skip to main content

e2e-template-testing

End-to-end validation of coordinator and agent template changes

Quellinformationen

Repository
github/gh-aw
Letzte Quellaktivität
17. August 2026 um 21:20
Erkannte Sprache von SKILL.md
Englisch
Sterne
5.200
Forks
564

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
e2e-template-testing
description
End-to-end validation of coordinator and agent template changes
domain
development
confidence
high
source
manual
## Context Squad's coordinator prompt (`squad.agent.md`) and agent charters (e.g. `scribe-charter.md`) are shipped as templates in `.squad-templates/`. Changes to these files affect how every squad session behaves — but unit tests can't catch prompt-level regressions because the prompts are interpreted by an LLM at runtime. This skill describes how to validate template changes end-to-end by running real squad sessions against a locally-built CLI that includes your modified templates. ## When To Use - You changed `.squad-templates/squad.agent.md` (coordinator prompt) - You changed `.squad-templates/scribe-charter.md` or other agent charters - You changed `.squad-templates/notes-protocol.md` or helper scripts - You added new conditional blocks (e.g. state-backend-aware spawn templates) - You modified the init scaffolding that writes templates to target repos ## Prerequisites - **Node.js** ≥20, **npm** ≥10 - **Git** CLI - **GitHub Copilot CLI** (`copilot` or `ghcs`) installed - A local clone of the squad repo on your feature branch ## Workflow ### Step 0 — Post initial tracking comment (**FIRST action — before anything else**) If `PR_NUMBER` and `REPO` are both set, **the absolute first thing you do** — before fast-fail checks, before building, before creating any repos — is post the initial tracking comment with all steps marked as `:hourglass_flowing_sand: Pending`. This gives reviewers immediate visibility that a run is in progress and what to expect. ```powershell $runStart = Get-Date $body = @" ## E2E Progress - PR $env:PR_NUMBER | Step | Status | Started | Duration | |---|---|---|---| | 1. Fast-fail checks (build :cd: link :cd: ``squad version``) | :hourglass_flowing_sand: Pending | --:-- | -- | | 2. Create test repo(s) | :hourglass_flowing_sand: Pending | --:-- | -- | | 3. ``squad init`` + file verification | :hourglass_flowing_sand: Pending | --:-- | -- | | 4. Run sessions | :hourglass_flowing_sand: Pending | --:-- | -- | | 5. Verify outcomes | :hourglass_flowing_sand: Pending | --:-- | -- | | 6. Record verdicts + post final comment | :hourglass_flowing_sand: Pending | --:-- | -- | | Symbol | Meaning | |---|---| | :hourglass_flowing_sand: | Not started | | :arrows_counterclockwise: | Running | | :white_check_mark: | Passed | | :x: | Failed | | :warning: | Passed with caveats | *Run started: $($runStart.ToString('HH:mm')) — all steps pending* "@ $tmpFile = [System.IO.Path]::GetTempFileName() $utf8NoBom = New-Object System.Text.UTF8Encoding $false [System.IO.File]::WriteAllText($tmpFile, $body, $utf8NoBom) $response = gh api "repos/$env:REPO/issues/$env:PR_NUMBER/comments" --method POST --field "body=@$tmpFile" | ConvertFrom-Json $env:COMMENT_ID = $response.id Remove-Item $tmpFile -Force Write-Host "Progress comment posted — ID: $($response.id)" ``` After posting, immediately update the comment to mark Step 1 as `:arrows_counterclockwise: Running` (do NOT wait — this is a two-step sequence: post all-pending, then immediately update to Step 1 running). Then proceed to Step 1. See **Progress Reporting** for the full comment lifecycle and update patterns. **⚠️ STOP before continuing:** If posting the comment fails (network error, auth error), abort the run and report the failure. Do not proceed silently without a tracking comment. ### Step 1 — Build the CLI from your branch ```bash cd /path/to/squad # your feature branch npm install npm run build -w packages/squad-sdk && npm run build -w packages/squad-cli # Link so `squad` command uses your local build (workspace flag — no cd required) npm link -w packages/squad-cli ``` Verify: `squad version` output includes the `-preview` suffix (e.g., `x.y.z-preview`), confirming the local dev build is active. If the output shows a plain semver without `-preview`, the globally-installed npm package is still in use — re-check the link step. See [CONTRIBUTING.md — Making the `squad` Command Use Your Local Build](../../../CONTRIBUTING.md) for the full guidance on local dev versioning. ### Step 2 — Create a disposable test repo ```bash mkdir /tmp/sq-test-1 && cd /tmp/sq-test-1 git init echo "# Test Project" > README.md echo '{"name":"test-project","version":"1.0.0"}' > package.json mkdir src echo "export function hello() { return 'world' }" > src/index.ts git add -A && git commit -m "init: test project" ``` Keep the project small — you only need enough for the coordinator to recognize a codebase and hire a team. ### Step 3 — Init a squad with your modified templates ```bash squad init # If testing a specific feature (e.g. state backends): # squad init --state-backend git-notes ``` Verify the init produced the expected files: ```bash ls -la .squad/ cat .squad/team.md # should have ## Members with 3+ agents cat .squad/config.json # should reflect any CLI flags you passed ``` ### Step 4 — Run a real session and capture output Use the Copilot CLI's `-p` flag with `--allow-all-tools` for non-interactive sessions. `--allow-all-tools` is **required** for automated/non-interactive runs — without it, tool calls (including file writes) prompt for confirmation and block. ```powershell # PowerShell (Windows) copilot --agent squad --allow-all-tools -p "Picard, decide what testing framework to use. Write your decision." ` 2>&1 | Tee-Object evidence/session-task.log ``` ```bash # Bash (macOS/Linux) copilot --agent squad --allow-all-tools -p "Picard, decide what testing framework to use. Write your decision." \ 2>&1 | tee evidence/session-task.log ``` Alternatively, set the `COPILOT_ALLOW_ALL=1` environment variable instead of the flag. For multi-turn workflows, run sequential sessions: ```powershell # Session A: give the team a task copilot --agent squad --allow-all-tools -p "prompt A" 2>&1 | Tee-Object evidence/session-A.log # Session B: verify state persisted copilot --agent squad --allow-all-tools -p "What decisions has the team made?" 2>&1 | Tee-Object evidence/session-B.log ``` ### Step 5 — Verify the outcome Check that your template change had the expected effect. Common checks: ```bash # State location (for state-backend changes) git notes --ref=squad list # git-notes backend git ls-tree -r squad-state # orphan backend ls .squad/agents/*/history.md # worktree backend # Coordinator behavior (grep session log) grep "STATE_BACKEND" evidence/session-task.log grep "spawn" evidence/session-task.log # File tree diff git diff --stat HEAD~1 # what changed on working branch git log --all --oneline # commits across all branches ``` ### Step 6 — Record the verdict Create an `evidence/verdict.md` in each test repo: ```markdown ## Test: [scenario name] **Backend:** worktree | git-notes | orphan | two-layer **Branch:** [your feature branch] **Result:** PASS | PARTIAL | FAIL **Duration:** Xm Ys ### What was verified - [ ] Coordinator identified feature correctly (from session log) - [ ] Agent was spawned via `task` tool (not simulated) - [ ] team.md has ## Members with 3+ agents - [ ] State landed in correct location - [ ] No unexpected side effects ### Evidence files - session-task.log — full session output - git-log.txt — `git log --all --oneline` ### Notes [anything unusual or noteworthy] ``` Record the wall-clock time from the start of Step 1 (fast-fail checks) to the end of Step 6 (verdict posted). This is the full E2E run duration for this scenario. ## Progress Reporting Use this section only when you are running E2E validation for an open PR. If `PR_NUMBER` and `REPO` are both set, post and maintain a live tracking comment in the PR thread. If either value is missing (for example, a local-only run), skip progress reporting silently. ### Start the tracking comment (Step 0 — see Workflow above) The initial comment must be posted as **Step 0** — the absolute first action before anything else. See the Step 0 block in the Workflow section for the exact code. The subsections below describe how to **update** the comment at each step boundary. For reference, here is the initial all-pending comment body posted in Step 0: 1. Post a PR comment before Step 1 begins: ```bash gh pr comment "$PR_NUMBER" --repo "$REPO" --body "## E2E Progress\n\n| Step | Status | Started | Duration | |---|---|---|---| | 1. Fast-fail checks (build · link · \\`squad version\\`) | ⏳ Pending | --:-- | -- | | 2. Create test repo(s) | ⏳ Pending | --:-- | -- | | 3. \\`squad init\\` + file verification | ⏳ Pending | --:-- | -- | | 4. Run sessions | ⏳ Pending | --:-- | -- | | 5. Verify outcomes | ⏳ Pending | --:-- | -- | | 6. Record verdicts + post final comment | ⏳ Pending | --:-- | -- | \n| Symbol | Meaning | |---|---| | ⏳ | Not started | | 🔄 | Running | | ✅ | Passed | | ❌ | Failed | | ⚠️ | Passed with caveats |" ``` 2. Capture the comment ID immediately after posting it: ```bash COMMENT_ID=$(gh api "repos/$REPO/issues/$PR_NUMBER/comments" --jq '.[-1].id') ``` 3. Treat Step 1 as in progress as soon as the comment exists. Update the body so Step 1 shows `🔄 Running` and every later step remains `⏳ Pending`. ### Update the tracking comment after every step boundary 1. When marking a step `🔄 Running`, record `$startTime = Get-Date` and store the `HH:MM` start time in that row's **Started** column. 2. Edit the existing comment in place; do not post a new progress comment: ```bash gh api --method PATCH "repos/$REPO/issues/comments/$COMMENT_ID" --field body="..." ``` 3. When marking a step `✅`, `❌`, or `⚠️`, compute `$duration = (Get-Date) - $startTime` and format it as `"{0}m {1}s" -f [int]$duration.TotalMinutes, $duration.Seconds`. 4. Update the completed step row to `✅`, `❌`, or `⚠️`, keep its original `HH:MM` value in **Started**, and write the formatted duration in **Duration**. 5. Keep all previously completed rows unchanged. 6. Mark the next step as `🔄 Running` and set its **Started** value. 7. Leave later steps as `⏳ Pending` with `--:--` for **Started** and `--` for **Duration**. 8. If a step fails and you stop early, still update the comment so the failed step shows `❌` with its original start time and computed duration, and Step 6 becomes `🔄 Running` while you prepare the final verdict. ### Use this status legend in the comment | Symbol | Meaning | |---|---| | ⏳ | Not started | | 🔄 | Running | | ✅ | Passed | | ❌ | Failed | | ⚠️ | Passed with caveats | ### Use exact step names and order Keep these six rows in this exact order every time you update the comment: 1. Fast-fail checks (build · link · `squad version`) 2. Create test repo(s) 3. `squad init` + file verification 4. Run sessions 5. Verify outcomes 6. Record verdicts + post final comment ### Handle Windows comment bodies safely On Windows PowerShell 5.1, use the **`--field body=@file`** pattern to post comment bodies. Write the content to a temp file using UTF-8 **without BOM**, then pass `--field "body=@$tmpFile"` to `gh api`. This is more reliable than piping JSON through `--input -` on PS 5.1, which can silently corrupt multi-byte characters even with `[Console]::OutputEncoding = UTF8`. Key rules: - Use `New-Object System.Text.UTF8Encoding $false` (the `$false` disables the BOM). `[System.Text.Encoding]::UTF8` writes a BOM which GitHub renders as a stray character (``) at the start of the comment. - Use `--field "body=@$tmpFile"`, NOT `--input -` or `--input filename`, for comment body updates. The `@` prefix tells `gh` to read the field value from the file rather than treating the path as a literal string. - Clean up the temp file after posting. - Scrub any local absolute paths from the body before posting (see PII Protection section). ```powershell $step1StartTime = Get-Date $step1Started = $step1StartTime.ToString('HH:mm') $step1Duration = (Get-Date) - $step1StartTime $step1DurationText = "{0}m {1}s" -f [int]$step1Duration.TotalMinutes, $step1Duration.Seconds $step2StartTime = Get-Date $step2Started = $step2StartTime.ToString('HH:mm') $body = @" ## E2E Progress | Step | Status | Started | Duration | |---|---|---|---| | 1. Fast-fail checks (build · link · `squad version`) | :white_check_mark: Passed | $step1Started | $step1DurationText | | 2. Create test repo(s) | :arrows_counterclockwise: Running | $step2Started | -- | | 3. `squad init` + file verification | :hourglass_flowing_sand: Pending | --:-- | -- | | 4. Run sessions | :hourglass_flowing_sand: Pending | --:-- | -- | | 5. Verify outcomes | :hourglass_flowing_sand: Pending | --:-- | -- | | 6. Record verdicts + post final comment | :hourglass_flowing_sand: Pending | --:-- | -- | | Symbol | Meaning |
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen