Skip to main content

evidence-analysis

Evaluates collected data to determine if evidence is sufficient to decide, then frames the outcome as CONTINUE, PAUSE, ROLLBACK CANDIDATE, or INCONCLUSIVE. Activate when triggered by CF-06 or CF-07 from the release-decision framework, or when user says "analyze results", "should I ship this", "continue or rollback", "is this significant", "what do the results say", "has it been long enough". Do not use when data collection has not started.

Aller à l'installation

Informations de source

Dépôt
featbit/featbit-release-decision-agent
Dernière activité de la source
5 mai 2026 à 08:48
Langue détectée de SKILL.md
anglais
Étoiles
1
Forks
0

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
3 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
evidence-analysis
description
Evaluates collected data to determine if evidence is sufficient to decide, then frames the outcome as CONTINUE, PAUSE, ROLLBACK CANDIDATE, or INCONCLUSIVE. Activate when triggered by CF-06 or CF-07 from the release-decision framework, or when user says "analyze results", "should I ship this", "continue or rollback", "is this significant", "what do the results say", "has it been long enough". Do not use when data collection has not started.
license
Apache-2.0
metadata
{"author":"FeatBit","version":"1.1.0","category":"release-management"}
# Evidence Analysis This skill handles **CF-06: Evidence Sufficiency** and **CF-07: Decision Framing** from the release-decision framework. CF-06 and CF-07 are handled together because they represent a continuous decision: first determine if evidence is sufficient, then frame what the evidence says. ## When to Activate - Data is being collected and the user wants to know whether to decide now - The user is impatient to interpret weak or early evidence - Results exist and a go/no-go decision is needed - Project stage is `measuring` or `deciding` ## On Entry — Read Current State Before doing any work, read the project from the database using the `project-sync` skill's `get-experiment` command. Check these fields: | Field | Purpose | |---|---| | `entryMode` | `"expert"` → user pre-filled setup + possibly data via the wizard; do not ask them to re-describe the experiment | | `primaryMetric` | The metric that decides the outcome | | `guardrails` | Metrics that must not degrade | | `hypothesis` | The causal claim being tested (may be empty in expert mode — don't block on it) | | `stage` | Current lifecycle position | | `experimentRuns[*].inputData` | JSON observed-data snapshot pasted via the wizard, shape `{metrics:{event:{variant:{n,k}|{n,sum,sum_squares},inverse?}}}` | | `experimentRuns[*].analysisResult` | Output of `runAnalysis` / `runBanditAnalysis`; may already exist | - If `primaryMetric` is empty AND `entryMode !== "expert"` → redirect to `measurement-design`. In expert mode, the primary metric lives in `experimentRuns[*].primaryMetricEvent` even if the top-level `primaryMetric` text field is blank. - If `stage` is `deciding` → a decision may already exist; check experiment records before re-analyzing - If experiment records already have a `decision` field → may only need to review, not re-decide ### Pulling observed data When `experimentRuns[*].inputData` is populated, that JSON *is* the observed data — you do not need track-service, ClickHouse, or live event queries. Parse it directly and use it for analysis. Trigger analysis by POSTing to `/api/experiments/<experimentId>/analyze` with `{runId}`. The endpoint automatically falls back to the stored `inputData` when `featbitEnvId` / `flagKey` are not wired up (expert-mode experiments with no FeatBit flag). The response includes `dataSource: "live" | "stored"` so you can tell the user where numbers came from. If the user asks "do you have my data?" or "can you see what I entered?", read `inputData` and confirm concretely: event name, per-variant n/k (or n/sum/sum_squares), guardrail events, inverse flags — not "I can't reach the database." ### Respect the `inverse` flag — do not override the verdict The analyzer's `verdict` and `p_harm` values **already account for each metric's `inverse` flag** (read from the guardrails JSON on the experiment and from `metrics[event].inverse` in `inputData`). They are authoritative. **Hard rules:** 1. **Do not quote metrics that don't exist.** The analyzer outputs `p_harm` (probability of harm) and `p_win` (probability of win) — they are complements. Never write things like "P(win) ≈ 0% so rollback" when the actual output field is `p_harm = 0`. That's inventing numbers. 2. **Do not override the analyzer's `verdict` silently.** If the analyzer says `"guardrail healthy"` and `p_harm = 0`, that is the evidence. Your job is to frame it, not flip it. 3. **When your intuition disagrees with the verdict, flag the configuration.** If a guardrail shows `rel_delta` with large magnitude (say |Δ| ≥ 50%) but `verdict: "guardrail healthy"` and `p_harm ≈ 0`, the most likely explanation is that `inverse` is set the wrong way for what the user actually meant. Ask, don't assume: > "gtest moved from 2.3% to 20% (+770%), and the analyzer reports P(harm)=0 with verdict `healthy`. That's because the guardrail is configured as 'higher is better' (`inverse=false`). If this metric is actually 'lower is better' (e.g. error rate, abandonment, latency), flip `inverse` in the setup and re-run — the verdict will change. Which did you mean?" 4. **If the user confirms the config is correct**, go with the analyzer's verdict. A +770% move on a higher-is-better guardrail is not harm. 5. **If the user confirms inverse was wrong**, they need to toggle it in the wizard (Edit setup) and re-analyze. Do not pretend the flipped-direction numbers apply to the current run record — they don't until the re-analysis writes a fresh `analysisResult`. This rule exists because a previous run produced a ROLLBACK decision by misquoting `p_harm=0` as `P(win)≈0%` and ignoring `inverse=false`. That fabrication is not allowed. ## Decision Actions ### Evidence sufficiency check (CF-06 first) Before interpreting results, confirm: 1. **Simultaneous?** — Are both variants measured over the same time window? 2. **Sufficient volume?** — Sample per variant ≥ `minimumSample` in the experiment record. If below this floor, the Gaussian approximation is unreliable — do not interpret P(win) or risk values yet. 3. **Risk has had a chance to converge?** — Read the experiment's `analysisResult` and check that `risk[trt]` and `risk[ctrl]` are not both still very high (> 0.02). If both are high, the posterior is still wide — more data is needed regardless of what P(win) shows. 4. **Clean window?** — Were there external events (promotions, outages, holidays) that could contaminate the data? 5. **Instrumentation verified?** — Are events firing correctly for both variants? 6. **SRM check passed?** — `analysisResult` includes a χ² SRM check. If it flags an imbalance (p < 0.01), do not interpret metric results until the traffic split issue is resolved. If any check fails, the right move is NOT to decide — it is to wait, fix, or extend. ### Decision framing (CF-07) Once evidence is sufficient, read the experiment's `analysisResult` and frame the outcome using exactly one of these categories: - **CONTINUE** — Primary metric P(win) ≥ 95% and risk[trt] is low. Guardrail P(win) all > 20%. Proceed with planned expansion. - **PAUSE** — Primary metric P(win) 80–95%, or a guardrail P(win) ≤ 20%, or SRM check failed. Signal exists but is not clean enough to expand. Investigate before proceeding. - **ROLLBACK CANDIDATE** — A guardrail P(win) ≤ 5%, or primary metric P(win) ≤ 5%. Evidence of harm. Flag should be reverted. - **INCONCLUSIVE** — Sample below validity floor, or risk[trt] and risk[ctrl] both still high, or primary metric P(win) 20–80% after a full observation window. Extend window or revisit instrumentation. See [references/decision-framing-guide.md](references/decision-framing-guide.md) for how to write each category's decision statement and what counts as "low" for risk values. ### Produce the decision artifact Write a structured decision statement with: - The recommendation category - The evidence that supports it (numbers, not vague descriptions) - The link back to the original hypothesis - The explicit next action ## Operating Rules - Do not let urgency substitute for evidence - "Not enough data" is a valid and honest decision frame — do not dress it up when the real issue is impatience - Separate "we don't know yet" from "we know it's harmful" - Hand off to `learning-capture` immediately after the decision is made ### Persist State Use `Skill("project-sync", ...)` to sync state. Stage stays at `measuring` — no stage advance here (the project stage advances to `learning` only when `learning-capture` completes): ```python assert Skill("project-sync", f'update-state {experiment_id} --lastAction "Decision: {category}"').ok # stage stays at measuring — do NOT call set-stage here assert Skill("project-sync", f'record-decision {experiment_id} {slug} --decision {category} --decisionSummary "{summary}" --decisionReason "{reason}"').ok assert Skill("project-sync", f'decide-run {experiment_id} {slug}').ok assert Skill("project-sync", f'add-activity {experiment_id} --type decision_recorded --title "Decision: {category}"').ok ``` ## Execution Procedure ```python def analyze_evidence(project_id, user_message): state = Skill("project-sync", f"get-experiment {project_id}") if state.primaryMetric in ("", None): Skill("measurement-design", project_id) return active_run = pick_active_run(state) # run in collecting or analyzing status # --- 6-check sufficiency gate --- checks = [ check_simultaneous(active_run), check_volume(active_run), # n >= minimumSample per variant check_risk_convergence(active_run), check_clean_window(active_run), check_instrumentation(active_run), check_srm(active_run), # chi-sq p >= 0.01 ] if any(check.failed for check in checks): say(format_insufficiency(checks)) return # do not produce a decision; do not write record-decision # --- 6-rule classification cascade --- category = classify(active_run.analysisResult) # ROLLBACK: guardrail P(win) <= 5% or primary P(win) <= 5% # PAUSE guardrail: guardrail P(win) <= 20% # CONTINUE: primary P(win) >= 95% and risk[trt] low and all guardrails > 20% # PAUSE primary: primary P(win) 80-95% # INCONCLUSIVE: P(win) 20-80% after full window, or risk both still high # lean-control: P(win) < 20% but above ROLLBACK threshold summary, reason = build_decision_artifact(category, active_run) assert Skill("project-sync", f'update-state {project_id} --lastAction "Decision: {category}"').ok assert Skill("project-sync", f'record-decision {project_id} {active_run.slug} --decision {category} --decisionSummary "{summary}" --decisionReason "{reason}"').ok assert Skill("project-sync", f'decide-run {project_id} {active_run.slug}').ok assert Skill("project-sync", f'add-activity {project_id} --type decision_recorded --title "Decision: {category}"').ok Skill("learning-capture", project_id) ``` ## Signal Inference | Check | Rule | |---|---| | `primaryMetric` empty | Redirect to `measurement-design` | | No active run | Check experiment records — may need `experiment-workspace` to start one | | SRM check fails | Stop; do not interpret metric results; investigate traffic split | | Both risk values still high | More data needed; do not decide — wait | | User impatient with sample below floor | Explain: below `minimumSample`, Gaussian approximation is unreliable | | INCONCLUSIVE | Still requires a written decision artifact — "we don't know yet" is a valid and complete frame | ## Reference Files - [references/decision-framing-guide.md](references/decision-framing-guide.md) — CONTINUE/PAUSE/ROLLBACK CANDIDATE/INCONCLUSIVE language, decision statement template, common framing mistakes - [references/tool-featbit-abtesting.md](references/tool-featbit-abtesting.md) — FeatBit experiment dashboard, reading per-variant results, confidence interpretation
Voir sur GitHub