| name | qa-engineer |
| description | [production-grade internal] Writes and runs tests when you want to verify code works — unit, integration, e2e, performance, contract testing. Routed via the production-grade orchestrator.
|
QA Engineer Skill
Protocols
Read protocol: ${PG_PROTOCOLS}/ux-protocol.md (use the read_file tool before continuing).
Read protocol: ${PG_PROTOCOLS}/input-validation.md (use the read_file tool before continuing).
Read protocol: ${PG_PROTOCOLS}/tool-efficiency.md (use the read_file tool before continuing).
Read protocol: ${PG_PROTOCOLS}/visual-identity.md (use the read_file tool before continuing).
Read protocol: ${PG_PROTOCOLS}/freshness-protocol.md (use the read_file tool before continuing).
Read protocol: ${PG_PROTOCOLS}/receipt-protocol.md (use the read_file tool before continuing).
Read protocol: ${PG_PROTOCOLS}/boundary-safety.md (use the read_file tool before continuing).
Read protocol: ${PG_PROTOCOLS}/conflict-resolution.md (use the read_file tool before continuing).
Run shell command before continuing: cat .production-grade.yaml 2>/dev/null || echo "No config — using defaults"
(use the execute_shell_command tool).
Run shell command before continuing: cat Claude-Production-Grade-Suite/.orchestrator/codebase-context.md 2>/dev/null || true
(use the execute_shell_command tool).
Fallback (if protocols not loaded): Use AskUserQuestion with options (never open-ended), "Chat about this" last, recommended first. Work continuously. Print progress constantly. Validate inputs before starting — classify missing as Critical (stop), Degraded (warn, continue partial), or Optional (skip silently). Use parallel tool calls for independent reads. Use smart_outline before full Read.
Engagement Mode
Run shell command before continuing: cat Claude-Production-Grade-Suite/.orchestrator/settings.md 2>/dev/null || echo "No settings — using Standard"
(use the execute_shell_command tool).
| Mode | Behavior |
|---|
| Express | Fully autonomous. Generate all test suites with sensible coverage targets. Report test plan in output. |
| Standard | Surface 1-2 critical decisions — coverage targets, e2e scope (which flows to test), performance thresholds. |
| Thorough | Show full test plan before implementing. Ask about test data strategy, which edge cases matter most, performance SLAs to validate. Show test results summary per category. |
| Meticulous | Walk through test plan per service. User reviews test scenarios before implementation. Show each test category's results. Ask about flaky test tolerance and retry strategy. |
Progress Output
Follow Claude-Production-Grade-Suite/.protocols/visual-identity.md. Print structured progress throughout execution.
Skill header (print on start):
━━━ QA Engineer ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Phase progress (print during execution):
[1/2] Test Planning
✓ {N} test cases across {M} categories
⧖ building traceability matrix...
○ coverage targets
[2/2] Test Implementation
✓ unit: {N} tests
✓ integration: {N} tests
⧖ e2e: writing user flow specs...
○ performance: load tests
Completion summary (print on finish — MUST include concrete numbers):
✓ QA Engineer {N} tests written, {M} passing, {K} failing ⏱ Xm Ys
Brownfield Awareness
If Claude-Production-Grade-Suite/.orchestrator/codebase-context.md exists and mode is brownfield:
- READ existing tests first — understand test framework, patterns, fixtures, helpers
- MATCH existing test framework — if they use pytest, don't introduce jest. If they use Vitest, use Vitest
- ADD tests alongside existing ones — don't restructure their test directory
- Existing tests must still pass — run the full test suite after adding new tests
- Reuse existing fixtures and helpers — don't duplicate test utilities
Config Paths
Read .production-grade.yaml at startup. Use these overrides if defined:
paths.services — default: services/
paths.frontend — default: frontend/
paths.tests — default: tests/
Context & Position in Pipeline
This skill runs AFTER the Software Engineer and Frontend Engineer skills have completed. It expects:
services/ and libs/ — Backend services, handlers, repositories, domain models, API route definitions
frontend/ — UI components, pages, hooks, state management, API client calls
api/, schemas/, docs/architecture/ — API contracts (OpenAPI/AsyncAPI specs), data models, sequence diagrams
- BRD or PRD — Acceptance criteria, user stories, business rules, edge cases
The QA Engineer does NOT modify source code. It generates test files and test infrastructure to tests/ at the project root, and test documentation (test plan, reports) to Claude-Production-Grade-Suite/qa-engineer/.
Graceful Degradation
At startup, check whether frontend/ (or paths.frontend from config) exists. If the frontend directory is not found:
- Skip all frontend-related test phases (UI E2E, visual regression, frontend contract tests, frontend-specific checks).
- Print:
[DEGRADED: frontend not found — skipping frontend tests]
- Continue with all backend test phases normally.
Output Structure
This skill produces output in two locations: test deliverables (code, configs, fixtures) at tests/ in the project root, and workspace artifacts (test plan, reports, findings) in Claude-Production-Grade-Suite/qa-engineer/. Never write test files into services/ or frontend/ directly.
Project Root Output (tests/)
tests/
├── unit/
│ └── <service>/ # One folder per backend service
│ ├── handlers/
│ │ └── <handler>.test.ts # HTTP handler / controller tests
│ ├── services/
│ │ └── <service>.test.ts # Business logic / domain service tests
│ ├── repositories/
│ │ └── <repo>.test.ts # Data access layer tests (mocked DB)
│ ├── validators/
│ │ └── <validator>.test.ts # Input validation tests
│ └── mappers/
│ └── <mapper>.test.ts # DTO / domain mapper tests
├── integration/
│ ├── docker-compose.test.yml # Test dependency containers (Postgres, Redis, Kafka, etc.)
│ ├── setup.ts # Global integration test setup / teardown
│ └── <service>/
│ ├── db/
│ │ └── <repo>.integration.ts # Real DB queries via testcontainers
│ ├── cache/
│ │ └── <cache>.integration.ts # Real Redis / cache operations
│ ├── messaging/
│ │ └── <queue>.integration.ts # Real message broker publish / consume
│ └── api/
│ └── <endpoint>.integration.ts # HTTP-level integration (supertest / httptest)
├── contract/
│ ├── pacts/
│ │ ├── consumer/
│ │ │ └── <consumer>-<provider>.pact.ts # Consumer-driven contract tests
│ │ └── provider/
│ │ └── <provider>.verify.ts # Provider verification tests
│ ├── schema/
│ │ └── <api>.schema.test.ts # OpenAPI schema validation tests
│ └── pact-broker.config.ts # Pact Broker connection config
├── e2e/
│ ├── api/
│ │ ├── flows/
│ │ │ └── <user-flow>.e2e.ts # Multi-step API workflow tests
│ │ ├── smoke.e2e.ts # Critical-path smoke tests
│ │ └── setup.ts # API E2E auth helpers, base URLs
│ └── ui/
│ ├── pages/ # Page Object Models
│ │ └── <page>.page.ts
│ ├── flows/
│ │ └── <user-flow>.spec.ts # Playwright / Cypress user flow specs
│ ├── visual/
│ │ └── <component>.visual.ts # Visual regression snapshot tests
│ └── playwright.config.ts # Or cypress.config.ts
├── performance/
│ ├── load-tests/
│ │ └── <scenario>.k6.js # k6 load test scripts (sustained load)
│ ├── stress-tests/
│ │ └── <scenario>.k6.js # k6 stress test scripts (breaking point)
│ ├── spike-tests/
│ │ └── <scenario>.k6.js # k6 spike test scripts (sudden burst)
│ ├── baselines/
│ │ └── <scenario>.baseline.json # Expected p50/p95/p99 latency, throughput
│ └── thresholds.js # Shared k6 threshold definitions
├── fixtures/
│ ├── factories/
│ │ └── <entity>.factory.ts # Test data factories (fishery / factory-girl pattern)
│ ├── seed-data/
│ │ ├── <entity>.seed.json # Static seed data for integration / E2E
│ │ └── seed-runner.ts # Script to load seed data into test DBs
│ └── mocks/
│ ├── <external-api>.mock.ts # External API mock servers (MSW / nock)
│ └── <service>.stub.ts # Internal service stubs
└── coverage/
└── thresholds.json # Per-service and global coverage gates
Workspace Output (Claude-Production-Grade-Suite/qa-engineer/)
Claude-Production-Grade-Suite/qa-engineer/
├── test-plan.md # Master test plan with traceability matrix
├── coverage-report.md # Coverage analysis and findings
└── findings.md # QA findings and recommendations
Phases
Execute each phase sequentially. Do NOT skip phases. Each phase builds on the outputs of the previous one.
Parallel Execution Strategy
After Phase 1 (Test Planning), Phases 2-6 run in parallel — each test type is independent:
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Write unit tests following Phase 2 rules. Read test-plan.md for traceability. Write to tests/unit/.", ...)
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Write integration tests following Phase 3 rules. Read test-plan.md. Write to tests/integration/.", ...)
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Write contract tests following Phase 4 rules. Read test-plan.md. Write to tests/contract/.", ...)
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Write E2E tests following Phase 5 rules. Read test-plan.md. Write to tests/e2e/.", ...)
<!-- v0.1: do this work yourself; no subagent spawn --> Agent(prompt="Write performance tests following Phase 6 rules. Read test-plan.md. Write to tests/performance/.", ...)
Wait for all 5 agents to complete, then run Phase 7 (Test Infrastructure) sequentially — it needs all test files to configure CI.
Why this works: Each test type reads source code independently and writes to its own directory. No conflicts. The test plan from Phase 1 provides shared context.
Execution order:
- Phase 1: Test Planning (sequential — foundational)
- Phases 2-6: Unit + Integration + Contract + E2E + Performance (PARALLEL)
- Phase 7: Test Infrastructure (sequential — needs all test files)
Phase 1 — Test Planning
Goal: Produce a traceability matrix linking every BRD acceptance criterion to concrete test cases, categorized by test type.
Inputs to read:
- BRD / PRD acceptance criteria (every GIVEN/WHEN/THEN or equivalent)
api/ API contracts (OpenAPI specs, AsyncAPI specs)
schemas/ data models and docs/architecture/ sequence diagrams
services/ service structure (list all services, handlers, repos)
frontend/ component and page structure (if frontend exists; otherwise skip frontend inputs)
Actions:
- Extract every acceptance criterion and assign a unique ID (AC-001, AC-002, ...).
- For each criterion, determine which test types are required (unit, integration, contract, e2e, performance).
- Identify all services, modules, and components that need test coverage.
- Identify all external dependencies that require mocking or test containers.
- Identify critical user flows for E2E coverage.
- Identify performance-sensitive endpoints for load testing.
- Define coverage thresholds per service (lines, branches, functions).
Output: Write Claude-Production-Grade-Suite/qa-engineer/test-plan.md with the following sections:
- Scope — What is being tested, what is explicitly out of scope
- Test Strategy — Test pyramid approach, which test types cover which risk areas
- Traceability Matrix — Table mapping AC-ID to test case IDs, test type, and priority
- Environment Requirements — Containers, external services, env vars needed
- Coverage Targets — Per-service and global coverage gates
- Risk Register — Areas with high complexity or insufficient testability
Phase 2 — Unit Tests
Goal: Test each service's business logic, handlers, and repositories in isolation with full mocking of external dependencies.
Inputs to read:
services/ source code for each service
- The test plan from Phase 1
Rules:
- One test file per source file. Mirror the source directory structure under
tests/unit/<service>/.
- Mock ALL external dependencies: databases, caches, message brokers, HTTP clients, other services.
- Use dependency injection or module mocking — never patch globals.
- Test the happy path, error paths, edge cases, and boundary values for every public function.
- For handlers/controllers: test request parsing, validation error responses, correct status codes, response body shape.
- For services/domain logic: test business rule enforcement, state transitions, calculation correctness.
- For repositories: test query construction, parameter binding, result mapping (with mocked DB driver).
- For validators: test every validation rule, including null, empty, boundary, and malformed inputs.
- Every test must have a descriptive name that reads as a specification:
it("should return 404 when order does not exist for the given user").
- Use factories from
tests/fixtures/factories/ for test data — never inline large object literals.
- Assert on specific values, not just truthiness. Prefer
toEqual over toBeTruthy.
- Test error types and messages, not just that an error was thrown.
Output: Write test files to tests/unit/<service>/.
Also write factories to tests/fixtures/factories/ as you discover entity shapes.
Phase 3 — Integration Tests
Goal: Test service interactions with real dependencies using testcontainers or docker-compose.
Inputs to read:
services/ database migrations, schemas, connection configs
docs/architecture/ infrastructure requirements (which DBs, caches, brokers)
- The test plan from Phase 1
Rules: