| name | user-uat |
| description | Run an already-clear UAT block FOR the operator — execute each step, capture the real output, and auto-judge only the mechanically-checkable ones (exit code, output match, refusal text, DB/HTTP/file/log); escalate every judgment call with evidence. Removes the run-command → paste-output → relay loop. With --ui, also drives + vision-judges the visual-tier steps via /judge-ui instead of escalating them. Use when the operator has a concrete UAT/manual-smoke block (a plan M-step, a "commands + what to look for" table, or ad-hoc "run these, expect these") and wants the mechanical tier done for them. Invoke as "/user-uat [source] [--deep] [--ui] [--dry-run] [--yes-side-effects]". |
| user-invocable | true |
User UAT
Judging doctrine: the mechanical-gates-first, evidence-on-every-verdict, and low-confidence-escalates invariants behind this skill's partition live in _shared/judge-core.md — this skill instantiates them for the UAT-execution case.
Execute an already-clear UAT block so the operator doesn't have to be the mechanical
relay (run command → eyeball output → paste it back). The skill runs each step, captures
the real result, and auto-judges only the deterministically-checkable steps; everything
that needs judgment is escalated with the evidence attached.
It does two things deliberately NOT: it does not refine a fuzzy script (that is
/review-uat), and it does not replace the operator for checks that genuinely need a
human. The win is the mechanical tier — which is most of the volume.
When to use / not
- Use: a concrete UAT exists — a
Type: operator plan M-step, a build-phase "Manual UAT"
bundle, the "commands + what to look for" table from a handoff, or the operator pasting
"run these and tell me what happens." Steps have commands and observable expectations.
- Don't use to write or refine a UAT. If steps are ambiguous (verb with no object, "expect
X" with no observable, no pass criteria), STOP and delegate to
/review-uat — do not
guess what a step means. A run built on a guessed expectation is worse than no run.
Invocation
/user-uat # use the UAT block from the current conversation
/user-uat path/to/plan.md#M4 # run the M-step at this anchor
/user-uat --dry-run # classify + print what WOULD run (mechanical / side-effectful / escalated); run nothing
/user-uat --deep # also agent-JUDGE the judgment-class checks; each flagged 'agent-judged: <verdict> — confirm?'
/user-uat --ui # drive + vision-JUDGE the visual-tier steps via /judge-ui instead of escalating them to you
/user-uat --yes-side-effects # auto-run side-effectful steps too (trusted flow); default gates them
The partition (the load-bearing rule)
Operator UAT exists to catch what agent self-checks miss (agents grading agent-written work
codify regressions — toybox G2; the audit-wire-shape rule). So never auto-PASS a check that
isn't deterministic. Classify every step's action and verify (separately) into:
| Tier | The verify is… | Default behavior |
|---|
| Mechanical | exit code, "stdout contains X", "refuses with Y", a DB row / count, HTTP status, file on disk, a log line | Auto-judge PASS/FAIL — show the observed value as evidence |
| Agent-judgeable | "output looks grounded", "the ship-vs-park call was right", a diff reads sensibly | Escalate with evidence (default). With --deep: agent assesses too, labeled agent-judged: <verdict> — confirm? |
Visual (needs --ui) | a rendered screen state visible in one browser frame — layout, a component, copy, a list/table, "the right screen showed" | Without --ui: escalate (Human). With --ui: drive + vision-judge via /judge-ui — read-back-cross-checked; UNCERTAIN → escalate, never auto-PASS |
| Human | animation / motion, audio / sfx, real-device input, kid-facing feel, anything credentialed or physical the agent can't drive | Always escalate — a crisp one-line ask, never a verdict |
When classification is ambiguous, treat the verify as human and escalate. A
mechanical-judgement applied to a judgment-class check is exactly how the blind spot leaks
back in. Even with --ui, motion / audio / feel stay Human — a screenshot is one frame; it
can't see a countdown tick or hear a sound.
Flow
- Ground + classify (no guessing). For each step emit a classification line:
Step N (source: file:line / M-anchor) — action: <command>; verify: <expectation>; Tier: Mechanical / Agent-judgeable / Human
The source citation must appear on the per-step classification line, not just in a header.
Valid tiers are Mechanical, Agent-judgeable, Human — plus, only when --ui is
passed, Visual (a vision-judgeable subset carved out of Human: a rendered screen state, not
motion/audio/feel). There is no "Ungroundable" category. If a step can't be grounded or is
ambiguous → stop, report it, and point at /review-uat.
- Safety gate. Tag each command read-only/preview vs side-effectful (mutates state,
is outward-facing, or is hard to reverse — e.g. a real
goblin do that auto-ships into a
sibling repo, a deploy, an external send, a DB write/drop, a git push, starting a process
that writes or sends anything). Auto-run the safe ones; pause and confirm before each
side-effectful one (unless --yes-side-effects). Never rationalize a side-effectful step
as "probably read-only" to skip the confirmation gate. If reversibility / outward-facing-ness
is unclear, treat it as side-effectful and confirm (fail safe). Prefer the step's own
--dry-run/preview when it has one.
- Run the auto-tier. For each step, run its action (subject to the step-2 side-effect
gate — the action may be agent-run, or a human/side-effectful one you've confirmed), capture
stdout/stderr + exit code, then auto-judge only the mechanical-tier verify against its
concrete expectation. Show the actual observed value inline — data, not editorializing. A
mechanical FAIL stops the run (don't barrel past a failure into dependent steps), then
still emit the step-5 report for the steps that ran + which step failed (observed-vs-expected).
Long-running actions (a server start, a watcher): run them in the background, then poll
a readiness probe (health endpoint, listening port, or an expected log line) before running
any dependent verify — never block the run on a foreground server. A probe that never comes up
within its budget is a mechanical AUTO-FAIL with the captured log tail as evidence, and
stops the run like any other mechanical FAIL.
- Judgment tier. For agent-judgeable / human steps: present the captured evidence + the
expectation and escalate (default). With , also give the agent's assessment for
the agent-judgeable ones — flagged , with any uncertainty named.
Human-tier steps always escalate regardless of .
without , escalate like Human. With , delegate to
— it drives the screen, captures stage screenshots, and renders a vision verdict
; a corroborated PASS lands in the step-5 report
with its evidence (screenshot + read-back value), and an /low-confidence/pixels-vs-
read-back-disagree result falls back to the section. Use the project adapter for
bring-up + auth (e.g. toybox ); never drive an app instance you don't own.
Safety + discipline
- Never auto-PASS a non-mechanical check. Mechanical = auto; agent-judgeable = escalate
(or
--deep + label); human = always escalate.
- Gate side effects. Read-only/dry-run/preview auto-run; destructive or outward-facing
steps confirm first.
--yes-side-effects only for a flow the operator has declared trusted.
- Delegate, don't guess. Fuzzy/ungroundable script →
/review-uat, not a guessed run.
- Push back on state mismatch. If the operator (or
--deep) calls something PASS but the
mechanical check disagrees, surface the discrepancy verbatim (observed X, expected Y) plus
ONE disambiguating question — don't rubber-stamp (per feedback_uat_pushback_on_state_mismatch).
- Show the data. Every verdict cites the observed value; a verdict with no evidence is a
defect.
--ui: never drive an app instance you don't own. A vision flow that logs in with a UAT
PIN against the operator's real (or a parallel session's) app can lock out their account. Use
the project adapter's isolated bring-up, and check port ownership before driving.
Relationship to other skills
/review-uat — the refinement partner. It tightens a fuzzy UAT, and its --exec delegates
execution of the refined script HERE — user-uat is /review-uat --exec's execution target, as
well as the terse run-an-already-clear-one path that hands fuzzy input back to it. Refine with
review-uat, then run with user-uat.
/judge-ui — the visual-tier executor --ui delegates to: drives a browser flow,
captures stage screenshots, and renders a vision verdict cross-checked against a read-back
(UNCERTAIN → Human fallback). Project adapters (e.g. toybox /uat-ui) supply its bring-up +
auth. Pairs with --deep the way judge-ui pairs with this skill's Human-fallback.
/verify — runs the app to confirm a code change works; user-uat runs a defined UAT
script, partitioning auto-vs-human across its steps.
/build-phase, /build-step — their Type: operator / "Manual UAT" outputs are exactly
the blocks user-uat is built to execute.
/user-walkthrough, /user-shakedown — the two operator-acceptance siblings for a
just-built feature with no clear script yet. /user-uat EXECUTES an already-clear script;
/user-walkthrough is operator-DRIVEN exploration (you drive, the agent answers from source /
fixes small / logs big); /user-shakedown AUTONOMOUSLY CLOSES the resulting UAT ledger to zero
open items. Poke a fresh build with a walkthrough/shakedown; run a defined block with user-uat.