| name | fix-flaky-tests |
| description | Diagnose and fix flaky/nondeterministic tests in the Glific frontend (Vitest + React Testing Library + Apollo MockedProvider, plus Cypress e2e). Reproduce intermittency, classify the root cause (real-timer/polling races, cross-file timer leakage, unflushed act() updates, global/localStorage state leakage, MockedProvider mock exhaustion, Cypress infra flakes), apply the smallest safe fix (usually fake timers or proper isolation), and verify with repeated isolated + one clean concurrent run. Use when the user reports flaky tests or shares CI logs for intermittent failures. |
| disable-model-invocation | true |
| effort | xhigh |
Fix Flaky Tests (Frontend)
Evidence-first workflow for flaky Vitest unit/component tests and Cypress e2e. Do not guess — reproduce, classify with evidence, fix minimally, verify.
Required inputs
Do not begin until the user provides one:
- A list of flaky tests (file path, optionally test name), or
- CI logs showing the intermittent failure (the
Unit Testing / E2E Tests workflow log).
If neither is given, ask for it. Then extract the exact failing test name + file + error shape from the log — CI stderr is full of Apollo addTypename warnings; grep them out (grep -vE "An error occurred|apollo.dev") and look for the real FAIL … > <test name> / Failed Tests line.
Reproduce locally (targeted)
Vitest:
yarn test:no-watch src/path/to/File.test.tsx
yarn test:no-watch src/path/to/File.test.tsx -t "test name"
yarn test:no-watch src/containers/<area>
Two stress modes — run them one at a time:
- Isolated: loop the single file ~20–25×.
- Concurrent: loop the whole directory ~6× (parallel workers — this reproduces the CI event-loop-starvation condition that exposes timer races).
⚠️ Never run two stress loops at once. Saturating the CPU makes even healthy synchronous tests miss an act() flush and fail — a measurement artifact, not a real flake. One loop at a time, or you'll chase ghosts.
If it won't reproduce locally even concurrently, the flake is almost always real-timer/polling (see below) — fix it structurally; you don't need a local repro to justify fake timers.
Classify the root cause (with evidence)
-
Real-timer / polling races (most common here)
-
Cross-file timer / state leakage
- Symptoms: a sync test fails only when run with the rest of its directory.
- Fixes: fake timers (above) make leaked timers inert;
vi.clearAllMocks() + localStorage.removeItem(...) in beforeEach; ensure the component clears its timers/polling on unmount.
-
Unflushed state updates
Apply the smallest safe fix
- Prefer test/setup fixes when the product behavior is correct (it usually is for timing flakes).
- Change app code only with proven behavioral regression.
- Keep the same assertions; change only how the test reaches the state.
Hard guardrail
If you cannot demonstrate the root cause with evidence and the structural fix (fake timers / isolation) doesn’t apply, say so plainly — do not weaken assertions, add arbitrary waitFor timeouts, or skip/retry the test to hide it.
Verify (mandatory)
- Run the fixed file isolated ~20× → 0 failures.
- Run its directory once concurrently (clean, nothing else running) → pass.
npx tsc --noEmit → 0 errors; npx prettier --check <changed files> → clean.
- For Cypress: confirm a clean re-run of the shard.
Report: root cause + why it was flaky, files changed, the stress matrix (isolated N/N, concurrent pass), and tsc/prettier status.
Glific FE conventions to respect
- Vitest + React Testing Library; Apollo
MockedProvider for GraphQL; act from @testing-library/react.
- Use existing mocks under
src/mocks/; don’t hand-roll GraphQL responses when a mock exists.
- Keep
data-testids stable (tests and the codebase key on them).
- The Apollo
addTypename stderr noise is pre-existing — never treat it as the failure.
Related skills
- improve-code-coverage — when
codecov/patch fails after a fix.
- make-branch-ready-for-review — full green-CI workflow (includes re-triggering stuck runs and re-running Cypress infra flakes).