- name
- live-perf
- description
- Use when you want to measure and improve the performance of a feature in the running extension, hunt regressions, or perf-tune before shipping. Standalone for perf-tuning an existing feature; also invoked by /live-exercise Phase 7. Not for static code review without measurement.
# /live-perf — Live performance measurement and improvement
Exercise the feature to measure it, not to change it. Capture baselines. Evaluate measurements and audit GitLens conventions. Dispatch fix agents under three-tier discipline — never on speculation. Re-measure to confirm. Loop until no measured regressions and no convention violations remain.
**Measure-first or don't touch. Speculation goes to open-questions, never to agents.**
## When to use vs other skills
| Skill | Purpose | Mode |
| ---------------- | ------------------------------------------------------------ | ------------------- |
| `/live-inspect` | Primitive MCP tool reference | Primitive |
| `/live-exercise` | Audit-and-fix loop (functional, intent, polish, improvement) | Live + iterative |
| `/live-perf` | **Performance measurement + improvement loop** | **Live + measured** |
| `/review` | Standards + completeness checklist | Static |
| `/deep-review` | Code-path correctness tracing | Static |
Use `/live-perf`:
- **Standalone** when you want to perf-tune an existing feature without the full audit (render regressions, RPC chattiness, git-call storms, missing caching on hot paths)
- **Delegated from `/live-exercise` Phase 7** as part of the ship-gate convergence loop
Do not use `/live-perf`:
- For speculative optimization of code that isn't measured slow
- As a substitute for `/live-exercise` when functionality or intent are in question (fix those first; perf-tune what works)
- For static code review without any live measurement
## Prerequisites
- `vscode-inspector` MCP connected (auto-discovered via `.mcp.json`)
- Build currently passes (`pnpm run build:quick`)
- You can identify the **scope** (feature, lifecycle stage, user flow, diff, background op, or ad-hoc code path) — arbitrary "perf everything" invocations don't belong here
## Scope / Exercise / Baseline — derive from the invocation
Every invocation varies along three axes. Derive each from the user's request; ask when ambiguous. **The framework (three-tier discipline, convergence loop, exit) is invariant; only the axis values change.**
### Axis 1 — Scope (what to measure)
| Scope kind | Examples |
| ---------------- | ----------------------------------------------------------------- |
| Feature | Commit Graph scrolling, Home hydration, Timeline filters |
| Lifecycle stage | Extension activation, first-repo-open, post-authentication |
| User flow | Cold start → open repo → graph → click commit → inspect |
| Diff-derived | Changes on this branch (`git diff base...HEAD`), specific commits |
| Background op | Auto-fetch, background indexer, status-bar refresh, watchers |
| Ad-hoc code path | A specific method or entry point the user names |
### Axis 2 — Exercise strategy (how to drive it)
| Strategy | When to pick |
| --------------------- | -------------------------------------------------------------------------------------- |
| User-like interaction | Scope has visible UI affordances (default for features / flows) |
| Programmatic trigger | No UI affordance, or need exact timing (`execute_command`, event dispatch, `evaluate`) |
| Ambient observation | Background / cron-like operations — let it happen, observe over a time window |
| Cold launch | Lifecycle / activation scope — teardown + relaunch to force cold state |
| Stress variant | Scale / contention concern — rapid repeat, concurrent actions, large-data inputs |
### Axis 3 — Baseline strategy (what to compare against)
| Strategy | When to pick |
| ------------------- | ----------------------------------------------------------------------------- |
| Absolute threshold | Known budget (hydration ≤150ms, ≤3 git calls/action). Default when available. |
| Prior baseline | Regression hunt — compare to last captured `baseline.md` or to `main` |
| Comparative branch | Diff-derived scope — this branch vs base branch on same exercise |
| Averaged runs (3–5) | Stochastic metrics (wall clock, render timing). Default for timing. |
| Paired measurements | Scale questions — cold/warm, small/large-data, pre/post stress |
### Default axis combinations
| If the user says… | Scope | Exercise | Baseline |
| ------------------------ | --------------------- | ----------------------- | ------------------------------ |
| "startup time" | activation lifecycle | cold launch | absolute + averaged |
| "feature X" | feature | user-like | absolute + averaged |
| "this branch" | diff-derived features | user-like per feature | comparative vs base branch |
| "background fetch" | background op | ambient observation | absolute (CPU/IO budget) |
| "large repo / N commits" | feature + data shape | user-like on large repo | paired small-vs-large |
| "why is X slow" | user-named flow | programmatic or user | absolute + per-stage breakdown |
### If ambiguous
Ask the user one question:
> "What's the scope — a specific feature, a lifecycle stage (activation/startup), a cross-feature user flow, the diff on this branch, or a background operation?"
Then infer exercise + baseline from the scope choice and the defaults above.
## Measurement categories — what you measure (independent of scope)
Regardless of scope, every measurement pass covers these four categories. Skip only when clearly not applicable (document why if skipped).
> **Default when running on Opus: delegate the raw measurement capture to the `inspector-driver` subagent (Sonnet 5)** — `Agent({ subagent_type: "inspector-driver", model: "sonnet", prompt: <the exact probes: performance.now() deltas, updateComplete timings, notification counts, N samples> })`. Sonnet captured a disciplined `updateComplete` measurement at parity in a smoke test. **You keep the three-tier classification (Measured/Convention/Speculation) and all fix decisions** — the driver only returns numbers. Cold-launch/startup metrics and baseline↔post-fix comparisons: have the driver reuse the instance you launched, or run those cold samples yourself, per the reset rule below.
### 1. Render / webview perf
What to measure:
- Initial render + hydration time per webview (from `launch` / refresh to Lit `updateComplete`)
- Re-render frequency during state transitions (Lit component `updated()` counts)
- Layout thrash / forced reflows (Performance API `layout-shift`, long tasks)
- Animation / transition smoothness (composited vs main-thread, frame drops)
Tools: `evaluate_in_webview` with `performance.now()`, `PerformanceObserver`, `performance.getEntriesByType("measure")`, `document.querySelector('...').updateComplete` for Lit.
> **Multi-webview perf**: With 2+ webviews open, untargeted `evaluate_in_webview` silently runs in the first one — usually the wrong one. Run `list_webviews` first, then scope every measurement with `webview_url`/`webview_index` (full targeting reference in `/live-inspect`).
### 2. RPC / data transfer
What to measure:
- Notifications fired per user action (count by type)
- Payload size per notification (large arrays flagged)
- N+1 patterns (request loops instead of batched fetches)
- Round-trip count per flow
Tools: instrument the RPC layer via `evaluate_in_webview` to wrap `rpcController`/`rpcClient` under `src/webviews/apps/shared/rpc/`. `read_logs` for `RpcHost`/`RpcLogger` traffic in `src/system/rpc/logger.ts`. For extension host side, `read_logs` with pattern matching on the notification names.
### 3. Hot-path code audit
What to audit (code-level, not measured):
- Uncached repeated operations (multiple awaits to the same getter / provider method)
- Missing `@memoize` on pure, frequently-called getters
- Missing `GitCache` / `PromiseCache` usage on git-derived data
- Serialized `await`s where `Promise.all` fits
- Missing debounce / throttle on frequent events (scroll, resize, input, mousemove)
Audit the code under scope — a feature's source, the diff of a branch, the activation path, etc. Use Read and Grep.
### 4. Git calls
What to measure / audit (git cost is spawn-heavy, not execution-heavy):
- Redundant git commands across components (the same `git log`/`git status` fired multiple times for one user action)
- Unbatched parsing (one piece of info per command instead of parsing multiple from a single output)
- Missing cache use on repeat callers (same git result re-fetched)
- Git-triggered refresh storms (one external change triggering N refreshes)
Tools: `read_logs` with pattern matching on `git ` command logs (GitLens logs git calls with debug logging). Code audit for patterns in `packages/git-cli/src/providers/`. `evaluate` on the extension host to count command invocations over a user-action window.
## Three-tier discipline
Every finding is classified into one of three tiers. **The tier determines whether agents dispatch and under what rules.**
| Tier | Rule | Action | ID prefix |
| --------------- | ----------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | --------- |
| **Measured** | Quantified regression vs baseline OR a measurable hot path above threshold | **Dispatch fix agent.** Post-measurement required to confirm improvement. | `PR` |
| **Conventions** | Violation of a GitLens convention (missed cache/memoize on hot getter, serialized awaits, missing debounce) | **Dispatch liberty-fix agent.** Log entry in `decisions.md`. No baseline needed — convention IS the bug. | `PC` |
| **Speculation** | "This could probably be faster" — no measurement, no established convention | **File in `open-questions.md`. Do NOT touch.** | `PS` |
Rules:
- **Measure-first for Measured-tier.** Capture baseline BEFORE touching anything. Capture post-fix measurement AFTER. If post-fix doesn't improve vs baseline, revert and re-investigate — don't ship "fixes" that aren't fixes.
- **Convention-tier can dispatch on audit alone.** Missing `@memoize` on a repeated pure getter IS a bug — you don't need to measure the savings to justify the fix.
- **Speculation NEVER dispatches.** If you can't measure it and it's not a convention violation, it's an open question for the user.
## The loop
### 1. Derive scope / exercise / baseline
Pick an axis value for each of Scope, Exercise, and Baseline — from the user's invocation, the `/live-exercise` handoff, or the default combinations table above. If any axis is ambiguous, ask the user the single scope question above and infer the rest from defaults.
- When invoked from `/live-exercise` Phase 7: scope is whatever live-exercise was run against; exercise defaults to user-like; baseline defaults to absolute + averaged.
- When invoked standalone: derive from the user's request. "Startup," "feature X," "this branch," "background," "large repo" all map cleanly via the defaults table.
- Read `goals.md` if present for stated perf requirements (thresholds, budgets). These become Axis-3 targets.
Record the chosen axes at the top of `baseline.md` so the intent is auditable.
### 2. Baseline — drive the exercise (do not change code)
`launch` VS Code if not already running. For the chosen Exercise strategy:
1. Enable relevant logging / instrumentation (git debug log, RPC logger, PerformanceObserver).
2. **Drive the exercise according to Axis 2**:
- User-like: click/scroll/type through the flow, repeat per baseline strategy
- Programmatic: `execute_command`, event dispatch, or `evaluate` — repeat per baseline strategy
- Ambient: leave VS Code running, observe over a time window (e.g., 2–5 minutes), don't drive it
- Cold launch: teardown + relaunch for each run (startup mode needs cold state every time)
- Stress: rapid repeat, concurrent, or large-data variant per the baseline's paired-measurement need
3. Capture measurements for each category that applies:
- **Render**: `evaluate_in_webview` with `performance.measure`, `updateComplete` timing
- **RPC**: log-parsed notification counts + payload sizes
- **Git**: `read_logs` with git command pattern, counted over the exercise window
- **Hot-path**: N/A (audit only, not measured)
4. Record baseline numbers in `.work/live/<feature>-perf/baseline.md`:
```markdown
# <Feature> — Perf Baseline
Recorded: <ISO date>
## Render
- Home webview hydration: 220ms (avg 5 runs)
- Repeat render on branch switch: 85ms
## RPC
عرض على GitHub