[Planning] Use when calibrating estimates from actual code, diff, PR scope, and developer time.
disable-model-invocation
false
argument-hint
<plan-file> | --changes | --pr <number>
Quick Summary
Goal: Produce a 3-way estimation calibration report — pre_impl_estimate (from plan) vs true_estimate (from observed scope) vs actual_time (from git/user) — yielding two INDEPENDENT signals: developer execution variance and estimation model calibration variance.
Why two signals matter: They are confounded if not separated. If actual >> pre-impl, the bug could be (a) developer was slow, OR (b) the model under-estimated scope. Without computing TRUE from observed scope, you cannot tell which. Single-sample calibration has near-zero statistical power — the skill always reports this.
UI files (component/template/style) — count, group by screen
Backend handlers/entities/repos — count, classify per backend tier table
Tests (unit/integration/e2e) — count test files, count test cases (grep describe|it|Fact|Test\b)
Migrations / contracts / shared code — flag separately
Step 3: Run Blast Radius pass on observed scope
Touched files / components — count
Of those, complex (>500 LOC area, multi-handler, central) — count
Downstream consumers — use code graph trace if available: python .claude/scripts/code_graph trace <file> --direction both --json for changed entry-point files
Shared/common code touched — yes/no
Regression scope — list affected areas
Step 4: Apply canonical estimation framework to observed scope
Apply each tier table (UI / backend / test / risk margin / risk factors) from the inline framework below to the OBSERVED scope. Output:
true_likely_days (single midpoint)
true_min_days = likely × 0.9
true_max_days = likely × (1 + risk_margin)
true_estimate = '<min>-<max>d' range
Step 5: Get actual time
Try in order:
Git: timestamp of first commit on feature branch → timestamp of last commit (or merge commit). Convert to working days (8h business days, exclude weekends).
PR: open time → merge time. Same conversion.
Ask user via AskUserQuestion: "Git suggests N working days from first commit to merge. How much was actual coding time? (excludes meetings, code-review wait, context switches, vacations)"
ALWAYS surface the gap between elapsed time and reported coding time — they are different signals.
Risk factors — predicted vs applicable; note any new factors that surfaced (e.g., regression-fan-out not flagged but should have been)
Calibration suggestion — ONLY if user has run this skill ≥3 times with consistent direction. Single-sample → state "no statistical power, log this sample for future calibration"
Confidence — state confidence level for each verdict; uncertainty about actual time goes here
Step 9: Persist sample (optional)
If user wants longitudinal tracking, append the calibration row to plans/_estimation-samples.csv:
After ≥5 rows, run pattern detection on the CSV: if scope_var_pct is consistently negative (model over-estimates), suggest tier adjustment; if consistently positive (under-estimates), suggest adding risk factors or widening tier.
Estimation Framework (canonical — applied in Step 4)
The canonical framework lives in the Estimation Framework sync block at the end of this skill; Step 4 applies it verbatim to the observed (post-hoc) scope.
Output Report Template
# Estimation Calibration Report — <planorbranchname>## Summary
| Metric | Range / Value | Source |
| ----------------- | -------------------------- | ---------------------------------------- |
| Pre-impl estimate | <min>-<max>d (likely <m>d) | <planpathfrontmatter> |
| TRUE estimate | <min>-<max>d (likely <m>d) | observed scope (post-hoc) |
| Actual time | <n>d | git <firstcommit→merge>, user-confirmed |
**Scope variance** (TRUE vs pre-impl): <±n>% — <under/over/matched>**Execution variance** (actual vs TRUE likely): <±n>% — <fast/slow/matched>## Verdict
| Signal | Direction | Magnitude | Confidence |
| ------------------- | ----------------------------------------------- | --------- | ----------------- |
| Estimation model | <toooptimistic / toopessimistic / calibrated> | <±n>% | <low/medium/high> |
| Developer execution | <fast / slow / on-pace> | <±n>% | <low/medium/high> |
## Per-Layer Breakdown
| Layer | Predicted tier | Observed tier | Delta |
| ------------ | ------------------ | ------------------ | --------------- |
| UI | … | … | … |
| Backend | … | … | … |
| Tests | … cases | … cases | … |
| Blast radius | … areas, … complex | … areas, … complex | … |
| Risk factors | <predictedlist> | <applicablelist> | <added/removed> |
## Calibration Suggestions-<Ifsinglesample> No model adjustment from one data point. Logged to `plans/_estimation-samples.csv` (row N). Re-run /estimate-actual on future plans to build calibration corpus. Suggested adjustment after ≥3-5 samples with consistent direction.
-<Ifpatternacrosssamples> e.g. "UI tier 'Compose components into NEW screen' overshoots in 4/5 samples by ~0.5d → suggest splitting into two tiers OR widening band to 1-2.5d"
## Caveats- Actual time derived from <git/user>; <listanyuncertainty:weekends, code-reviewdays, vacationsexcluded?>- Pre-impl estimate format <range/single-point/missing> — comparison <exact/approximate>- Confidence in TRUE estimate: <high/medium/low> — observed scope <fullyvisible / partiallyobscured>
Anti-Rationalization Anchors
Evasion
Rebuttal
"Single sample is enough — clearly the dev was slow"
NO. Without separating scope from execution variance, you confound model error and performance. State signal + caveat.
"Use git timestamps as actual time"
Wrong. Includes weekends, meetings, code-review wait, sleep. Always confirm with user.
"Skip TRUE estimate — just compare pre-impl vs actual"
That's the data point that's MISSING and exactly why estimates don't improve over time. Never skip Step 4.
"Apply hindsight to pump up TRUE estimate"
Use the SAME framework that was used for pre-impl. Hindsight bias inflates TRUE and falsely vindicates the original estimate.
"One signal is fine, no need to split"
Two signals is the entire point. Performance review needs execution variance; model tuning needs scope variance. Confounded data is unactionable.
Estimation Framework — Bottom-up first; SP DERIVED; output min-max range when likely ≥3d. Stack-agnostic. Baseline: 3-5yr dev, 6 productive hrs/day. AI estimate assumes Claude Code + project context.
Method:
Blast Radius pass (below) — drives code AND test cost
Without tests, SP drops ≥1 bucket? → tests dominate; state explicitly
Reasoning called out UI vs backend vs blast vs risk factors? → if missing, add
AI Mistake Prevention — Failure modes to avoid on every task:
Re-read files after context changes. Context compaction, resume, or long-running work can make memory stale; verify current files before acting.
Verify generated content against source evidence. AI hallucinates APIs, names, claims, and document facts. Check the relevant source before documenting or referencing.
Check downstream references before deleting or renaming. Removing an artifact can stale docs, generated mirrors, configs, and callers; map references first.
Trace the full impact chain after edits. Changing a definition can miss derived outputs and consumers. Follow the affected chain before declaring done.
Verify ALL affected outputs, not just the first. One green check is not all green checks; validate every output surface the change can affect.
Assume existing values are intentional — ask WHY before changing. Before changing a constant, limit, flag, wording, or pattern, read nearby context and history.
Surface ambiguity before acting — don't pick silently. Multiple valid interpretations require an explicit question or stated assumption with risk.
Keep shared guidance role-relevant. Universal guidance must help every receiving skill or agent; code-specific obligations belong only in code-specific protocols.
Closing Reminders
Protocols in force (concise digest of the SYNC/shared blocks this skill carries):
AI Mistake Prevention: verify generated content against evidence, trace downstream references, verify all affected outputs, re-read after context loss, surface ambiguity.
Critical Thinking: Traced file:line proof per claim, confidence >80% to act, no guess-as-fact.
IMPORTANT MUST ATTENTION compute TRUE estimate using the SAME canonical framework — fair comparison requires identical methodology
IMPORTANT MUST ATTENTION separate developer execution signal from model calibration signal — never collapse to single verdict
IMPORTANT MUST ATTENTION never claim model adjustment from a single sample — explicitly state "needs ≥3 samples for signal"
IMPORTANT MUST ATTENTION never trust git timestamps as coding time — always ask user to confirm/override
IMPORTANT MUST ATTENTION list per-layer deltas (UI/backend/tests/blast) — aggregate variance hides where model went wrong
IMPORTANT MUST ATTENTION use min-max ranges for both pre-impl and TRUE — comparing single points is dishonest about uncertainty
IMPORTANT MUST ATTENTION apply Blast Radius pass on observed diff before applying tier tables
IMPORTANT MUST ATTENTION persist samples to plans/_estimation-samples.csv for longitudinal calibration
IMPORTANT MUST ATTENTION state confidence per verdict — uncertainty about actual time goes in caveats
[IMPORTANT] Use TaskCreate to break ALL work into small tasks BEFORE starting.
Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
MUST ATTENTION apply critical + sequential thinking — every claim needs appropriate traced evidence (file:line for repo/code claims; source URL or artifact section for research, product, content, and docs claims); confidence >80% to act, <60% DO NOT recommend. Anti-hallucination: never present guess as fact, admit uncertainty freely, cross-reference independently, stay skeptical of own confidence.
MUST ATTENTION apply AI mistake prevention — verify generated content against evidence, trace downstream references before deleting or renaming, verify all affected outputs, re-read files after context loss, and surface ambiguity before acting.
[TASK-PLANNING] Before acting, analyze task scope and systematically break it into small todo tasks and sub-tasks using TaskCreate.