Diff-aware AI browser testing โ reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP, ARIA-tree-first) with pass/fail reporting. Use when testing UI changes, verifying PRs before merge, or running regression checks on changed components.
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Claude Code 2.1.220+. Requires agent-browser >= 0.25.0 (Rust-native, no Playwright).
description
Diff-aware AI browser testing โ reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP, ARIA-tree-first) with pass/fail reporting. Use when testing UI changes, verifying PRs before merge, or running regression checks on changed components.
{"keywords":["expect","test my changes","browser test","diff test","test what I changed","test the UI","visual regression","check my changes"],"examples":["test my changes before I push","expect โ run browser tests on what I changed","test the login flow after my auth refactor","run visual regression on the dashboard"],"anti-triggers":["cover","unit test","generate tests","verify","implement","npm test"]}
["command -v agent-browser >/dev/null 2>&1 || echo 'Warning: agent-browser not installed โ run npm install -g agent-browser'"]
Expect โ Diff-Aware AI Browser Testing
Analyze git changes, generate targeted test plans, and execute them via AI-driven browser automation.
Note: If disableSkillShellExecution is enabled (CC 2.1.91), the agent-browser install check won't run. Verify it's installed: npx agent-browser --version.
/ork:expect # Auto-detect changes, test affected pages
/ork:expect -m "test the checkout flow"# Specific instruction
/ork:expect --flow login # Replay a saved test flow
/ork:expect --target branch # Test all changes on current branch vs main
/ork:expect -y # Skip plan review, run immediately
Core principle: Only test what changed. Git diff drives scope โ no wasted cycles on unaffected pages.
Argument Resolution
ARGS = "[-m <instruction>] [--target unstaged|branch|commit] [--flow <slug>] [-y]"# Parse from full argument stringimport re
raw = ""# Full argument string from CC
INSTRUCTION = None
TARGET = "unstaged"# Default: test unstaged changes
FLOW = None
SKIP_REVIEW = False# Extract -m "instruction"
m_match = re.search(r'-m\s+["\']([^"\']+)["\']|-m\s+(\S+)', raw)
if m_match:
INSTRUCTION = m_match.group(1) or m_match.group(2)
# Extract --target
t_match = re.search(r'--target\s+(unstaged|branch|commit)', raw)
if t_match:
TARGET = t_match.group(1)
# Extract --flow
f_match = re.search(r'--flow\s+(\S+)', raw)
if f_match:
FLOW = f_match.group()
raw.split():
SKIP_REVIEW =
1
# Extract -y
if
'-y'
in
True
STEP 0: MCP Probe + Prerequisite Check
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) โ probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")
# Verify agent-browser is available (Rust-native, no Playwright)
Bash("command -v agent-browser || npx agent-browser --version")
# If missing: "Install agent-browser: npm i -g agent-browser"# Load agent-browser's own self-serving skill/workflow docs (required since 0.25.x)
Bash("agent-browser skills get agent-browser")
CRITICAL: Task Management
# 1. Create main task IMMEDIATELY
TaskCreate(
subject="Expect: test changed code",
description="Diff-aware browser testing pipeline",
activeForm="Running diff-aware browser tests"
)
# 2. Create subtasks for each pipeline phase
TaskCreate(subject="Check fingerprint (skip if unchanged)", activeForm="Checking fingerprint") # id=2
TaskCreate(subject="Scan git diff and classify changes", activeForm="Scanning diff") # id=3
TaskCreate(subject="Map changes to routes/URLs", activeForm="Mapping routes") # id=4
TaskCreate(subject="Generate AI test plan", activeForm="Generating test plan") # id=5
TaskCreate(subject="Execute tests via agent-browser", activeForm="Executing browser tests") # id=6
TaskCreate(subject="Compile test report", activeForm="Compiling report") # id=7# 3. Set dependencies for sequential phases
TaskUpdate(taskId="3", addBlockedBy=["2"]) # Diff scan needs fingerprint check
TaskUpdate(taskId="4", addBlockedBy=["3"]) # Route map needs diff results
TaskUpdate(taskId="5", addBlockedBy=["4"]) # Test plan needs route map
TaskUpdate(taskId="6", addBlockedBy=["5"]) # Execution needs test plan
TaskUpdate(taskId="7", addBlockedBy=["6"]) # Report needs execution results# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress") # When starting
TaskUpdate(taskId="2", status="completed") # When done โ repeat for each subtask
Pipeline Overview
Git Diff โ Route Map โ Fingerprint Check โ Test Plan โ Execute โ Report
Phase
What
Output
Reference
1. Fingerprint
SHA-256 hash of changed files
Skip if unchanged since last run
references/fingerprint.md
2. Diff Scan
Parse git diff, classify changes
ChangesFor data (files, components, routes)
references/diff-scanner.md
3. Route Map
Map changed files to affected pages/URLs
Scoped page list
references/route-map.md
4. Test Plan
Generate AI test plan from diff + route map
Markdown test plan with steps
references/test-plan.md
5. Execute
Run test plan via agent-browser
Pass/fail per step, screenshots
references/execution.md
6. Report
Aggregate results, artifacts, exit code
Structured report + artifacts
references/report.md
Phase 1: Fingerprint Check
Check if the current changes have already been tested:
Read(".expect/fingerprints.json") # Previous run hashes# Compare SHA-256 of changed files against stored fingerprints# If match: "No changes since last test run. Use --force to re-run."# If no match or --force: continue to Phase 2
Floor is >= 0.25.0; current tested release is 0.33.1 (see upstream-version-tested). Commands below hold across this range. 0.30+ adds agent-browser read (agent-readable text extraction) and the --restore / --namespace session-restore workflow for stable, isolated browser state across agent runs. 0.33.0 adds agent-browser a11y [url], an embedded axe-core audit (WCAG tag filtering, selector scoping, iframe-aware text/JSON output) available as both a CLI command and an MCP tool. The commands documented below are unchanged from 0.32.x through 0.33.1.
Area
Command
Notes
Snapshot
agent-browser snapshot -i
ARIA tree w/ @eN refs. -C/--cursor was removed in 0.22
expect_task = Agent(
subagent_type="ork:expect-agent",
prompt=f"""Execute this test plan:
{test_plan}
For each step:
1. Navigate to the URL
2. Execute the test action
3. Take a screenshot on failure
4. Report PASS/FAIL with evidence
""",
run_in_background=True,
model="sonnet",
max_turns=50
)
# Stream agent-browser progress line-by-line instead of polling (CC 2.1.98+)# Each stdout line from agent-browser arrives as a notification โ useful for# catching a failing step early rather than waiting for the full plan.# Full pattern: Read("${CLAUDE_PLUGIN_ROOT}/skills/chain-patterns/references/monitor-patterns.md")
Monitor(pid=expect_task.agent_id)
# For long test plans (>3 min typical), notify on completion โ requires# Remote Control + "Push when Claude decides" config (CC 2.1.110+).# Skip silently if the user doesn't have Remote Control enabled.if test_plan_duration_estimate > 180:
PushNotification(
message=f"ork:expect complete โ {passed}/{total} steps passed on {len(affected_urls)} pages",
status="proactive"
)
When the dev stack is live (/ork:dev), saving any .tsx, .jsx, .css, or .scss file (and Next.js route files like app/**/page.tsx, pages/**/*.tsx) emits a nudge to run /ork:expect <route>. The hook (posttool/ui-change-detector) is default-on and:
skips silently if /ork:dev hasn't booted (no agent-browser session to attach to);
enforces a 30-second cooldown per route to prevent spam on rapid saves;
honors .claude/state/expect-skip.<sessionId> as a per-session opt-out (write any content);
honors ORK_EXPECT_AUTO=0 for an env-level kill switch.
Route resolution: app/dashboard/page.tsx โ /dashboard, pages/settings.tsx โ /settings, component / global-style edits โ / (home as proxy). Route groups like app/(marketing)/pricing/page.tsx strip to /pricing.
ARIA snapshot recording (M125 #6)
After a passing run, the posttool/expect/snapshot-recorder hook persists the captured ARIA tree to .claude/state/expect-snapshots/<route-slug>/<parent-commit>.json. Subsequent /ork:expect <route> --diff runs compare against the most recent prior snapshot for that route โ surfaces structural regressions (added/removed buttons, label changes, hierarchy shifts) without needing a baseline screenshot.
For the snapshot recorder to fire, the expect run output must contain RUN_COMPLETED|passed, ROUTE|<route>, and ARIA|<json-summary> tags. The agent-browser-driven flow already emits these.
Docs-only changes โ unless you want to verify docs site rendering
Quality Bar
Done means all of these hold:
Test-plan scope is derived from the git diff for the chosen --target โ no unaffected page appears in the plan.
Every changed file maps to at least one tested route (via .expect/config.yaml or inferred convention) OR is explicitly excluded as API-only, generated, or docs-only.
Fingerprint check runs first; a diff unchanged since the last run skips execution instead of re-testing.
Unless -y is passed, the plan is presented for review before any browser action runs.
Each executed step reports PASS or FAIL with evidence (screenshot on failure), and the report's pass/fail totals match the steps actually run.
The report's exit code is non-zero whenever any step failed.