Skip to main content

measurement-design

Defines the primary success metric, guardrails, and event instrumentation for a hypothesis before exposure begins. Activate when triggered by CF-05 from the release-decision framework, or when user asks "how do I measure this", "what metrics should I use", "what should I track", "how do I instrument this", "event design", "what is the primary metric". Do not use when metrics are already well-defined and instrumentation is complete.

Ir a la instalación

Datos de origen

Repositorio
featbit/featbit-release-decision-agent
Última actividad en el origen
5 de mayo de 2026 a las 08:48
Idioma detectado de SKILL.md
inglés
Estrellas
1
Forks
0

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Explorador de archivos
3 archivos

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
measurement-design
description
Defines the primary success metric, guardrails, and event instrumentation for a hypothesis before exposure begins. Activate when triggered by CF-05 from the release-decision framework, or when user asks "how do I measure this", "what metrics should I use", "what should I track", "how do I instrument this", "event design", "what is the primary metric". Do not use when metrics are already well-defined and instrumentation is complete.
license
Apache-2.0
metadata
{"author":"FeatBit","version":"1.1.0","category":"release-management"}
# Measurement Design This skill handles **CF-05: Measurement Discipline** from the release-decision framework. Its job is to ensure every hypothesis has exactly one primary metric, a small set of guardrails, and an event schema that can produce evidence for the decision. ## When to Activate - Hypothesis exists but no primary metric is defined - User names multiple competing success metrics - User mixes goals, proxies, and diagnostics together - Instrumentation does not exist for the desired outcome - Project `primaryMetric` field is empty or contains a list ## On Entry — Read Current State Before doing any work, read the project from the database using the `project-sync` skill's `get-experiment` command. Check these fields: | Field | Purpose | |---|---| | `hypothesis` | Confirms causal claim exists | | `primaryMetric` | Current metric definition (may be empty) | | `guardrails` | Existing guardrail metrics | | `stage` | Current lifecycle position | - If `hypothesis` is empty → redirect to `hypothesis-design` - If `primaryMetric` is already defined and instrumentation is confirmed → may not need this skill - If `guardrails` already exist → build on existing rather than overwriting ## Core Principle **One primary metric decides the outcome. Everything else is a guardrail or diagnostic.** If two metrics both "decide" success, the decision is not yet sharp enough. Return to `hypothesis-design`. ## Decision Actions ### Define the primary metric Ask: "If this experiment runs for 2 weeks and you had to make a go/no-go decision with ONE number, what is that number?" The answer is the primary metric. One only. Then capture it as a structured object — NOT a paragraph. The five fields are: | Field | Purpose | Example | |---|---|---| | `name` | Short human-readable label (shows in the web UI's metric table) | `"Signup conversion"` | | `event` | Instrumented event key emitted from code (snake_case, no spaces) | `"signup_completed"` | | `metricType` | `binary` for conversion (fires 0/1) or `continuous` for a value (revenue, latency, count per user) | `"binary"` | | `metricAgg` | `once` (max 1 per user, for funnel conversion), `count` (tally all occurrences), `sum` (add numeric payloads), or `average` (mean of values per user) | `"once"` | | `description` | One-sentence rationale — why this metric decides go/no-go | `"Proportion of visitors that sign up — directly measures the H1 change."` | If the user gives a vague metric name (e.g. "signup rate"), probe until the event key, metric type, and aggregation are concrete. Don't proceed with a half-defined metric. ### Define guardrails (2–3 maximum) Ask: "What other metrics would concern you if they degraded significantly, even if the primary metric improved?" Common guardrails: - Error rate / p99 latency for the candidate variant - User satisfaction score or support ticket volume - A downstream conversion step after the primary metric Each guardrail has the same five fields as the primary metric, **plus**: | Field | Purpose | Example | |---|---|---| | `direction` | `increase_bad` (e.g. error rate, abandonment) or `decrease_bad` (e.g. downstream retention) | `"increase_bad"` | ### Design the event For each metric, define: - Event name (what action fires it) - Required properties (user_key, session_id, relevant context) - Where in the user journey it fires - Whether it needs to be associated with a flag evaluation for experiment analysis ### Verify instrumentation completeness Check: can the current codebase emit this event? If not, instrumentation must be built before exposure begins. ## Operating Rules - Do not allow exposure to start without confirmed instrumentation - One primary metric only — push back on lists - Guardrails protect against harm, not success. They should not be optimized for. - When multiple experiments share a flag or user pool, verify traffic allocation strategy before calculating sample size. Sequential experiments get full traffic; concurrent experiments with mutual exclusion get only their slice. An underpowered experiment due to traffic splitting is worse than waiting for a sequential slot. See `reversible-exposure-control` → [multi-experiment-traffic.md](../reversible-exposure-control/references/multi-experiment-traffic.md) for patterns. - Hand off to `reversible-exposure-control` when instrumentation is complete - Hand off to `evidence-analysis` once data collection is underway ### Persist State Use `Skill("project-sync", ...)` to sync state. All three writes are required, and instrumentation must be confirmed before writing: ```bash PRIMARY='{"name":"Signup conversion","event":"signup_completed","metricType":"binary","metricAgg":"once","description":"Proportion of visitors that sign up."}' GUARDRAILS='[{"name":"Checkout abandonment","event":"checkout_abandoned","metricType":"binary","metricAgg":"once","direction":"increase_bad","description":"must not rise"}]' npx tsx $HOME/.claude/skills/project-sync/scripts/sync.ts update-state <experiment-id> \ --primaryMetric "$PRIMARY" \ --guardrails "$GUARDRAILS" \ --lastAction "Metrics defined" npx tsx $HOME/.claude/skills/project-sync/scripts/sync.ts set-stage <experiment-id> measuring npx tsx $HOME/.claude/skills/project-sync/scripts/sync.ts add-activity <experiment-id> --type stage_update --title "Metrics defined" ``` **`primaryMetric` must be a JSON object** with `{name, event, metricType, metricAgg, description?}`. The web UI renders each field as its own column (NAME / EVENT / TYPE / AGG) — do NOT jam a paragraph into `name`. Rationale goes in `description`. **`guardrails` must be a JSON array** of objects with the primary-metric shape plus `direction` (`increase_bad` or `decrease_bad`). One entry per guardrail, never a single string or newline-separated text. `sync.ts update-state` validates both fields' JSON shape and enums; if validation fails it will print what's missing and exit non-zero. ## Execution Procedure ```python def design_measurement(project_id, user_message): state = Skill("project-sync", f"get-experiment {project_id}") if state.hypothesis in ("", None): Skill("hypothesis-design", project_id) return # --- primary metric --- # ask: "if you had ONE number to make a go/no-go decision, what is it?" primary_metric = elicit_primary_metric(state, user_message) # --- guardrails (2–3 max) --- guardrails = elicit_guardrails(state) # --- event design --- events = design_events(primary_metric, guardrails) # --- instrumentation gate --- instrumentation_confirmed = confirm_instrumentation(events) if not instrumentation_confirmed: say("Instrumentation must be confirmed before exposure begins.") return # do not advance stage until confirmed primary_json = json.dumps({ "name": primary_metric.name, "event": primary_metric.event, "metricType": primary_metric.metric_type, # binary | continuous "metricAgg": primary_metric.metric_agg, # once | count | sum | average "description": primary_metric.rationale, }) guardrails_json = json.dumps([ { "name": g.name, "event": g.event, "metricType": g.metric_type, "metricAgg": g.metric_agg, "direction": g.direction, # increase_bad | decrease_bad "description": g.rationale, } for g in guardrails ]) assert Skill("project-sync", f'update-state {project_id} --primaryMetric {shlex.quote(primary_json)} --guardrails {shlex.quote(guardrails_json)} --lastAction "Metrics defined"').ok assert Skill("project-sync", f"set-stage {project_id} measuring").ok assert Skill("project-sync", f'add-activity {project_id} --type stage_update --title "Metrics defined"').ok Skill("reversible-exposure-control", project_id) ``` ## Signal Inference | Check | Rule | |---|---| | `hypothesis` empty | Redirect to `hypothesis-design` | | User lists multiple "primary" metrics | Push back — ask which ONE decides the go/no-go | | Guardrail count > 3 | Ask which 2–3 matter most; trim the rest to diagnostics | | Instrumentation not confirmed | Block stage advance; do not write `set-stage measuring` | | Traffic split planned (concurrent experiments) | Flag that sample size must be calculated on the reduced traffic slice, not full traffic | ## Reference Files - [references/event-schema-design.md](references/event-schema-design.md) — TrackPayload shape, event naming conventions, metric-to-event mapping, anti-patterns - [references/tool-featbit-sdk.md](references/tool-featbit-sdk.md) — FeatBit SDK track() usage, experiment event association, sendToExperiment
Ver en GitHub