| name | blind-calibration |
| description | Use when stating confidence ("I'm sure", "definitely", "99%"), giving confidence intervals, making probabilistic claims, or reviewing a track record of predictions. Applies overconfidence calibration research (Lichtenstein & Fischhoff; Alpert & Raiffa) — subjective certainty systematically exceeds accuracy. |
Calibration
Overview
Lichtenstein, Fischhoff & Phillips (1982; figures from Fischhoff, Slovic & Lichtenstein 1977) and decades of calibration studies: when people put the odds at 100:1, they're right ~73% of the time; "100% certain" answers are wrong 15–30% of the time. Alpert & Raiffa (1982): 98% confidence intervals capture the true value only ~60% of the time, and 50% interquartile ranges only ~33% — intervals would need to be roughly 2–3× wider (implied by the coverage gap, not reported as such). Confidence is a feeling generated by fluency and coherence of the story in your head (Kahneman: "confidence is a feeling, not a judgment"), and it tracks story quality, not truth.
The blind spot: certainty is experienced as a property of the world. It's a property of your narrative.
When to Use
- About to say "definitely", "certainly", "I'm 100% sure", "this can't fail"
- Producing ranges/intervals (estimates, forecasts, error bars)
- High confidence formed quickly, from fluent/familiar material
- Evaluating whether to double-check something that "doesn't need it"
The Procedure
- Translate the feeling into a bet: "would I stake something real at these odds?" 95% sure = comfortable losing 1 time in 20 at stakes that hurt. Usually the honest number drops immediately.
- Apply the standard correction: your "99%" historically behaves like ~73–90%. Treat your own extreme confidence as a trigger to verify, not a license to skip verification — 100% certainty is a flag, not a green light.
- Widen intervals until uncomfortable, then more: for ranges, state values you'd be surprised to see exceeded in either direction — genuine 90% intervals feel embarrassingly wide (2–3× your instinct).
- Audit the confidence source: is it verification (I ran it / read it / measured it) or fluency (it's familiar / coherent / I've said it before)? Only the first kind counts (see blind-dunning-kruger for domain-dependence).
- Keep score: log predictions with probabilities; review hit rates. Calibration is trainable (forecasters, meteorologists) but only against recorded outcomes — memory alone rewrites your record in your favor.
Quick Reference
| Utterance | Calibrated replacement |
|---|
| "This definitely works" | "Verified by X; unverified: Y" |
| "99% sure" (from memory) | "~85%, and it's cheap to check — checking" |
| "It'll take 2–3 days" | "2 days if nothing surprises, 5 realistic, 9 if X bites" |
| "That can't happen" | "I can't construct how it happens — which is different" |
Common Mistakes
- Blanket hedging everything ("might", "possibly") — that's underconfidence noise, not calibration. Numbers + named unverified parts beat vague hedges.
- Calibrating only when challenged. The correction applies most when no one is pushing back.
- Trusting remembered accuracy ("my estimates are usually right") — hindsight bias (Fischhoff 1975) edits your record; only written predictions count.