| name | lisa-exploratory-qa |
| description | First-time-user exploratory QA pass for ANY product type (DOM web app, HTTP/API backend, canvas game, CLI/library, IaC/CDK) that FEEDS THE LIFECYCLE. Use when asked to experience a product the way a brand-new end user would — driving its real consumer-facing interface via the `use-the-product` core (which detects the product type, resolves the per-environment mutation policy from .lisa.config.json so production data is never mutated by accident, and explores through the project's personas when it defines them) to find anything confusing, broken, or hard to use. Static route scans, HTTP fetches, screenshots alone, or console/network checks alone are not sufficient evidence. Instead of writing a report file, it files every finding as a tracked work item via lisa-tracker-write. A `ready` parameter controls whether those tickets are created build-ready (auto-picked-up by lisa-intake) or left in the backlog for human triage (default). For gaps in the automated test suite, use e2e-coverage-gaps instead. |
Exploratory QA
Overview
Experience the product the way a brand-new end user would: drive its real consumer-facing interface and actually try to use it, then surface anything confusing, broken, or hard to understand. This is a usability/experience pass, not a test-coverage audit (for that, use e2e-coverage-gaps). Every finding is filed as a tracked work item so it enters the Lisa lifecycle — no static report file.
How you drive the product is owned by the use-the-product core skill. Invoke it first: it detects the product type (web / API / game / CLI / IaC), resolves the target environment and its mutation policy (so you never mutate production without an explicit, justified opt-in), and discovers the project's personas so you can explore as each one. This skill supplies the QA lens — what to look for and how to file it.
Parameters
target-url | env (first positional) — what to explore (passed through to use-the-product).
ready=true|false — the build-ready state for the tickets this pass creates.
ready=true → created build-ready, so lisa-intake / the build-intake scanner auto-picks them up.
ready=false (default) → created in the backlog for a human to review and promote.
1. Set up
- Invoke
use-the-product to detect the product type, resolve the environment + mutation policy, and discover personas/subagents. Everything about how to drive the product, where, and how much you may mutate comes from there — do not re-derive it.
- For a DOM web app whose resolved non-production environment has mutation policy
full, check the optional Kane provider with lisa kane probe. If ready, invoke lisa-kane-browser for self-contained journeys and use its screenshots, HAR/network, console, and normalized verdict as evidence. Keep the run on a directly controlled browser for read-only/forbidden environments. Treat tool_failed/timed_out as provider findings, never as product bugs.
- Confirm the tracker is configured. Findings are filed as tickets, so read
tracker from .lisa.config.json (local overrides global). If unset, stop and report that the tracker must be configured (/lisa:setup:jira / :github / :linear) before exploratory QA can file findings — do not silently fall back to a report file.
- Read the
ready flag (default false).
2. Arrive cold and use it like a human
- Start with no prior knowledge — do not pre-read the codebase to learn the intended flows; discover them the way a user would. Form a first impression: is it obvious what this product is, what to do first, and where to go next?
- Then actually attempt real tasks through the interface (per
use-the-product's per-type playbook), exercising representative controls/endpoints/commands — a first-time user explores, makes mistakes, and tries the obvious thing.
3. The QA lens — what to look for
Cover at least these dimensions unless the user narrows scope. Most are universal; the web-specific ones are marked and have per-type equivalents in §4.
- Comprehension & labeling: user-facing copy must read like something a normal first-time user understands. Flag machine-style/developer labels (raw IDs, enum keys,
snake_case, null/undefined, untranslated i18n keys), admin/database terms ("metadata", "rows", "record", "entity"), unexplained jargon, unclear button/menu names, and meaningless icons. If a label would make a non-technical user ask "what does that mean?", file a clarity ticket.
- Data usefulness & context: facts, metrics, and tables must help a person understand the surface, not just prove extraction happened. If a user can't tell what a number or field refers to without rereading the raw source, file a ticket to add context (excerpts, labels, grouping, provenance) or hide it.
- Data volume & trust: compare visible data to what the surface promises. A rich view with only a few items can read as broken/filtered/still-loading. File a finding when sparse data isn't explained (result counts, active filters, reset affordances, coverage status, sample labeling).
- Controls & mental model: controls must match what they do (sort ≠ filter; search ≠ facet; finite domains want selects/typeaheads, not blank text inputs). Flag controls that make users guess spelling/casing, hide the option universe, or use the wrong component.
- Flow completeness & expected counterparts: a screen that gates access or shows one side of a standard paired flow must offer the other side, or a clear path to it. A new user must never hit a dead end with no next step. Flag: sign-in with no sign-up (or vice-versa), no account recovery ("Forgot password?", resend verification), no exit from a state (no sign-out; a modal/wizard with no back/close/cancel), one-way actions (create with no edit/delete where both are expected), unreachable entry points (a feature only reachable by guessing a URL; an empty state with no primary action). When the missing counterpart makes a core task impossible for a class of users, file a
Bug; otherwise a usability Improvement.
- Action preconditions & incomplete end-states: an action that needs multiple inputs or a prerequisite (compare, merge, bulk-edit) should guide the user to satisfy it — disable/explain until met, collect inputs first, or give the destination an in-place control. Actually trigger these and watch where they land; flag when a primary action fires under-satisfied and strands the user, lands on an empty/partial end-state with no next step, or is offered where it cannot succeed. Left unable to finish →
Bug; works but confusing → Improvement.
4. Type-specific depth
- DOM web app — breakpoints & layout integrity. Sweep a range of widths, including the in-between ones where clipping appears (e.g. 360, 390, 414, 600, 768, 834, 1024, 1280, 1440 plus ~900–1180 steps), and re-walk key paths at each. At every width, in addition to looking, take DOM measurements and treat as findings: container overflow (
documentElement.scrollWidth > clientWidth); clipped/offscreen controls (a control's getBoundingClientRect() falling outside the viewport or an overflow:hidden|clip|auto|scroll ancestor — e.g. a submit button cut off by its filter card); truncated meaningful text (scrollWidth > clientWidth / ellipsis on text that carries meaning); colliding controls (label overlapping an adjacent control with no gap). A primary/interactive control that is clipped, offscreen, or unreachable is a Bug, not an Improvement.
- HTTP / API backend. Judge the contract as a consumer experiences it: unclear/inconsistent status codes, error responses with no actionable message, payloads exposing internal shapes (raw enums, DB column names), missing pagination/among-counts, endpoints that 500 on obvious edge inputs. A broken/incorrect response is a
Bug; a confusing-but-working contract is an Improvement.
- Canvas game. Judge readability at gameplay scale, input responsiveness/latency, game-feel, unclear objectives, and silent state changes — not DOM breakpoints. A soft-lock or lost progress is a
Bug; unclear-but-playable friction is an Improvement.
- CLI / library. Judge help/output clarity, error messages, discoverability of commands, and surprising side effects. A wrong result / crash is a
Bug; confusing UX is an Improvement.
- IaC / CDK (read-only). From
cdk synth/diff: over-broad IAM, missing/opaque stack outputs, resources that don't serve their stated purpose, drift. A security-relevant misconfiguration is a Bug; unclear-but-correct infra is an Improvement.
5. Mutation
Whether you may create/edit/delete — and as which account — is set by the use-the-product mutation policy (read-only vs full, identity, production rules). Follow its Mutation Discipline when the policy is full (prefixed test data, identify + verify cleanup, record residue). Never mutate an env the policy marks read-only or forbidden; if a finding can only be confirmed by a forbidden mutation, file it as observed-and-blocked rather than escalating.
6. File findings as tracked work
No report file. Every finding becomes a leaf work item via lisa-tracker-write (the vendor-neutral writer — it dispatches to the configured tracker and runs the validation gate; never call a vendor *-write-* skill directly):
Kane never files the work item itself. Even when it marks a failure as a confirmed product bug, run Lisa's marker-based duplicate search, apply the QA lens, and route the finding through lisa-tracker-write exactly as for every other controller.
| Finding | issue_type | build_ready |
|---|
| User-visible bug (broken behavior) | Bug | the ready flag (default false) |
| Usability / UX / clarity issue | Improvement | the ready flag (default false) |
Each finding is a flat leaf, so build_ready applies directly — pass it explicitly on every create.
This skill is the named human-gate exception under ready-role-filing. Everywhere else in Lisa, a complete defect found during other work is filed with explicit build_ready: true so build-intake claims it next cycle. Exploratory QA is deliberately different, and that difference is ratified rather than drift: its findings are candidate defects and usability observations whose product significance is a human call, so auto-readying them would push judgment work into the build queue — precisely the failure the gate model exists to prevent. Its default ready=false is therefore an explicit human-gate marker, never a bare omission: every not-ready create passes human_gate: "exploratory finding — product significance is a human product call" alongside build_ready: false, and the writer stamps the auditable [lisa-human-gate] marker. An operator who wants a pass to feed the queue directly opts in with ready=true.
(The sibling e2e-coverage-gaps skills are the contrast: a missing automated test is not a product question, so they file build_ready: true by default.)
Each ticket MUST be a complete spec (the validator rejects thin tickets): a three-audience description; for a bug, exact reproduction steps, observed-vs-expected, the env / account / interface it occurred at, and evidence; for a usability issue, the observed friction, who it affects, where, and the proposed improvement; and Gherkin acceptance criteria for the fixed behavior.
Idempotency — don't spam duplicates
Re-running a pass must not refile the same finding. Before creating a ticket, search the tracker for a ticket carrying a stable marker [lisa-exploratory-qa] <finding-key> in its body (the <finding-key> is a stable slug of surface + symptom, e.g. settings-modal/horizontal-overflow@tablet). Per the rejection-detection rule's Proposal rejection memory section, that marker search MUST cover open AND closed tickets (with a body-enumeration fallback on search-index lag), and match by the marker, never by title. Then split on how any prior ticket closed:
- Open ticket carrying the marker → reference/update it instead; do not create a second.
- Closed as completed → does not suppress. A recurrence after a fix is a genuine regression, so file the finding.
- Closed as not planned (GitHub
stateReason == "not_planned"; the config-resolved won't-do/canceled equivalent on JIRA/Linear, including a Linear duplicate state) → a human declined this finding, so suppress it. Re-file only with evidence that postdates the decline, carrying BOTH the machine token (declined <date>; recurred <date> in <ref>) and a human acknowledgment sentence (You declined this on <date>. It has recurred (<date>, <ref>), so we're raising it once more for your review.).
Every filed finding ticket MUST end with the rejection-detection operator footer as a visible prose line so the operator knows which close-reason silences it:
To stop this from being raised again, close it as Not planned. Close it as Completed if it was fixed — a later recurrence may be re-filed as a regression.
Output
No report file. Emit a concise in-session summary:
- Scope: product type, target env + mutation level, persona(s) explored as, tool, build/version if visible, date.
- First impression: could a new user tell what the product is and what to do first?
- Findings filed, bucketed by type — each with its created or referenced ticket ref and build-ready state.
- Observed but not filed: anything noticed but intentionally not ticketed (including forbidden-mutation blocks), with why.
Run outcome
As the registered exploratory-bugs automation loop, this pass conforms to the
automation-runbook-contract rule: every invocation ends in exactly one of the six run outcomes
and records it, so a quiet run and a broken run are never mistaken for each other.
| This cycle's exit path | Run outcome |
|---|
Findings filed — one or more Bug / Improvement tickets created or referenced (§6) | candidate-proposed |
Clean pass — explored the personas and surfaces, nothing worth filing — or every candidate was suppressed by a prior decline (rejection-detection Proposal rejection memory): the summary MUST name the suppression count | nothing-needed |
Tracker unconfigured — the §1 stop path; findings cannot be filed — or the open-and-closed rejection-memory marker search could not run (tracker unreachable / credentials revoked): a memory check that could not run is a broken loop, never a silent nothing-needed | recovery-required |
The runbook's Retirement condition tripped — the trailing quiet window is empty AND this pass found nothing to file AND the project no longer ships an exploratory-qa surface — this row supersedes the nothing-needed row when it applies | policy-obsolete |
| A degradation that still let the pass explore (e.g. Kane unavailable, one persona unreachable) | the outcome it actually reached above, with the summary leading with the degradation — degradation never mints a seventh token |
Before invoking the run-record CLI, evaluate the Retirement condition first. If it applies,
select policy-obsolete as the sole outcome and do not record a prior nothing-needed result;
otherwise select the ordinary outcome from the table.
Record exactly one outcome per invocation through the run-record CLI, naming this loop's runbook
(the --summary is the operator-readable one-liner in the contract's exemplar voice — plain,
specific, actionable, e.g. Explored 4 personas; nothing confusing to file. — or, when a decline suppressed the only candidates, Explored 4 personas; 2 candidates suppressed by a prior decline — nothing new to propose. — for nothing-needed; and for a recovery-required from an unreadable decline check, Tracker unreachable during the decline check — restore credentials; nothing was filed this run.):
node "${CLAUDE_PLUGIN_ROOT}/scripts/automation-run-record.mjs" \
--loop-id exploratory-bugs --outcome candidate-proposed \
--summary "Explored 4 personas; filed 3 findings from the checkout flow — awaiting your flip to ready." \
--runbook .lisa/automations/exploratory-bugs.runbook.md [--ref <ticket-url>]...
If ${CLAUDE_PLUGIN_ROOT} is unset, resolve the plugin scripts directory directly — the built copy
plugins/lisa/scripts/automation-run-record.mjs or the source
plugins/src/base/scripts/automation-run-record.mjs. If recording still fails, degrade, never
abort (per automation-runbook-contract): note the recording failure in the run output and finish
the cycle — a recording failure is a degradation to report, never a reason to block the loop.
Retirement evaluation (every run). Evaluate this loop's runbook Retirement condition on
every run, exactly as the automation-runbook-contract rule's Retirement section defines it — this
skill conforms to that text and never restates or diverges from it. On top of the contract's two
conditions the runbook seeds a third domain conjunct — the project no longer ships an
exploratory-qa command surface to explore — which only tightens the test and never replaces it: a
quiet month on a product nobody broke is quality holding, not a reason to stop looking. Evaluate all
three. When all three hold, record policy-obsolete and file exactly ONE marker-deduped
teardown proposal through lisa-tracker-write (per tracked-work + integration-access-layer):
- Marker
<!-- [lisa-automation-retire] key=exploratory-bugs --> plus a visible prose line;
matched on the marker, never the title; searched open AND closed per rejection-detection's
Proposal rejection memory. Treat matches by close state: open suppresses another proposal;
Not planned suppresses another proposal unless new evidence postdates the rejection;
Completed means the prior approved action happened, so a later recurrence may be re-filed.
When an existing proposal suppresses filing, the run still records policy-obsolete and files
nothing — the outcome describes this run, while the ticket is filed exactly once.
- Labels
status:blocked + human-needed, carrying the contract's decision-ready packet. The
human-needed label marks the proposal human-owned: lisa-repair-intake recognizes it and never
re-dispatches it as stalled work.
- Evidence the date-filtered search result, this run's summary, the loop's current cadence
(the baseline an operator needs to choose a longer one), and a one-line summary of recent runs
read from
.lisa/automations/runs/exploratory-bugs.jsonl. Fill the rest of the packet the same
way every time: Work already attempted is the searches this run ran, and Risk of inaction is
that the loop keeps consuming schedule slots and tokens for nothing.
- How to answer names the three operator responses: approve — run
/lisa:tear-down-automations exploratory-bugs and only that loop registration goes away;
decline — close the proposal as
Not planned (closing it as Completed leaves a later re-file open) and the loop simply
continues; re-cadence — pick a longer cadence off that evidence and re-register with
/lisa:setup-automations instead of tearing down.
- Operator footer, verbatim, as on every loop-filed proposal (
rejection-detection):
To stop this from being raised again, close it as Not planned. Close it as Completed if it was fixed — a later recurrence may be re-filed as a regression.
The loop keeps running at its normal cadence until a human acts, and never deletes its own
registration.
Quality bar
- Explore as a true first-time user — judge clarity, not whether you (who can read the code) can figure it out.
- Every ticket must stand alone for an implementer who was not in the session.
- Do not claim cleanup succeeded unless verified.
- File per the
ready flag (default: backlog for human triage).
- Route automated-coverage gaps to
e2e-coverage-gaps; preserve unrelated repo changes.