| name | forecasting-judgment-foxes-and-track-record |
| description | Audits a forecaster's reasoning, track record, and process design for hedgehog overconfidence and formula-vs-intuition failures; invoke when reviewing forecasting judgment or credibility claims. |
| kind | skill |
| status | ready |
| provenance | {"principles":["P073","P078","P083","P085","P086","P087","P103","P121","P122","P152","P153","P154","P155","P156","P187","P188","P192","P197"],"claims":["C00713","C00714","C00716","C00717","C00718","C00719","C00720","C00721","C00722","C00724","C00749","C00750"],"evidence":[],"source_anchors":[],"authored_from_digest":"d2669973882e7b281ac8105ff44e4962a43eb81c6ade65b372028746e6291563"} |
Forecasting Judgment Foxes And Track Record
Purpose
This skill audits a forecaster's or analyst's reasoning style, claimed track record, and any
selection or accountability process behind a judgment under review: whether the reasoning shows
fox-like integrative complexity rather than hedgehog one-big-idea certainty, whether credit for a
"right" call survives the simpler explanation that a predisposition merely matched the outcome,
and whether fame, media prominence, or a bare "almost right" claim is being mistaken for a
credential. It checks that any claimed intuition or track record rests on an unflinching
postmortem and a genuinely regular, feedback-rich environment rather than on luck, and that the
forecasting and leading roles — along with any selection or accountability procedure attached to
the judgment — are structured to reward calibrated process over a decisive-looking outcome. This
skill audits the forecaster's process and credibility claim, not the domain content of the
forecast itself.
When to use
- A forecast, estimate, or claimed track record is offered as evidence of skill, and whether it
reflects fox-like integrative complexity or hedgehog one-big-idea certainty needs checking
(P192).
- A correct call is being credited to a forecaster's superior method, openness, or foresight, and
that credit needs testing against the simpler explanation that a prior predisposition happened
to match the outcome (P078).
- Multiple contributors, sources, or expert judgments are pooled into one estimate, or aggregated
into a claimed track record, and whether those inputs are independent and diverse — not merely
numerous or repeated — needs checking (P073, P152).
- A track record's credibility is being judged against an unstated or wrong baseline, or the
depth of domain knowledge it assumes has not been checked against a realistic floor of
usefulness (P153, P154).
- A forecaster's or team's credibility leans on public visibility, media demand, or a claim of
having been "almost right," and that claim needs testing rather than accepted at face value
(P155, P156).
- An expert's intuitive or gut judgment is being relied on, or a selection/prediction procedure in
an uncertain domain lets a human's global impression override a structured, formula-combined
score (P083, P187, P188).
- A forecast has just resolved, or a specific failure or setback is under review, and the
postmortem needs checking for hindsight bias, unexamined luck, or a closed-off rather than
data-driven response (P085, P121).
- A leadership decision blends forecasting language with command language, or a plan is moving
from deliberation into resolute execution, and whether the two modes — and the handoff between
them — were kept properly distinct needs checking (P086, P087).
- An accountability mechanism attached to a forecaster or analyst is under review, or a polarized
dispute could instead be converted into a concrete, benchmarked forecast the two sides agree
would settle it (P197, P122).
- The review must choose whether to model the reasoning under audit as a rational actor or invoke
a behavioral account, and only that narrow modeling choice — not how the actor values outcomes
or trades off gains and losses — is in question (P103).
Procedure
- Establish what is actually under review: the specific forecast, estimate, or track-record
claim, who produced it, and what evidence of "skill" or "credibility" is being asserted. Fix
this before judging the reasoning style or the record itself.
- Check the reasoning against fox/hedgehog markers: look for qualifying conjunctions ("however,"
"but") and visible efforts to weigh competing considerations against each other, rather than a
single confident narrative that never engages a rival account (P192). Where a forecaster is
credited with having "seen it coming," test that credit against the alternative that a
pre-existing predisposition simply matched how events unfolded — common to both correct and
mistaken forecasters, and no proof by itself of superior method or openness (P078).
- If confidence in the estimate rests on pooling multiple contributors, sources, or repeated
soundings, verify the inputs are genuinely independent and diverse before crediting the pool
with error-cancelling power (P073); verify any track-record claim rests on triangulated
evidence across many questions, experts, or cases, with confidence raised only as that
independent evidence converges, not from one dramatic hit (P152).
- Check the track record against the right baseline: during a volatile period the bar is a
naive or random baseline, not confident foresight, since breakpoints are far easier to spot
after the fact than at the time, and heightened expert confidence during a crisis is itself a
flag, not a credential (P153). Check that the floor of expertise the record assumes is
realistic — a thin-knowledge, case-by-case judgment is riskier than one built on a
well-informed generalist baseline, and additional specialist depth past that point buys little
(P154).
- Treat public visibility, media demand, or extreme displayed confidence attached to the
forecaster as a warning sign of overconfidence, not as a credential supporting the estimate
(P155). Where credit is claimed for being "almost right," discount that credit in proportion to
how self-servingly the defense is being invoked, especially where it surfaces only after a
miss (P156).
- Before accepting a claimed intuitive or gut judgment as trustworthy, check that the domain is
regular enough to be predictable and that the person has had prolonged practice with prompt,
valid feedback; in a zero-validity or "wicked" domain — sparse or misleading feedback — do not
accept the intuitive claim as insight (P083). Where the object under review is a selection or
prediction procedure in an uncertain domain, check that the final call is produced by combining
independently scored factors through a formula rather than a synthesized human impression
(P187); where an interview or evaluation procedure specifically is under review, check that it
scores roughly six independent traits one at a time on a fixed scale before combining by
formula, rather than a holistic read a halo effect can contaminate (P188).
- Check any postmortem behind the track record is unflinching: wins were examined as closely as
losses, hindsight bias was guarded against, offsetting errors or luck were ruled out before
crediting a good outcome to a good decision, and the reasoning is reconstructable from
contemporaneous notes rather than rebuilt after the fact (P085). Where a specific failure or
setback is under review, check the response reads as a try/analyze/adjust cycle that treats the
miss as data, not as an excuse that forecloses learning (P121).
Inputs
- The forecast, estimate, or track-record claim under review, including who produced it and what
claim of skill or credibility rests on it.
- The reasoning trail or write-up behind the judgment, so it can be checked for hedge language and
visible engagement with a rival account.
- Any postmortem notes or contemporaneous record from when the call was made, covering both wins
and losses.
- The forecaster's or team's performance across multiple questions or cases, where a track-record
or credibility claim is under review.
- What is known about the environment's regularity and feedback quality, where a claimed expert
intuition is under review.
- The design of any selection, hiring, or accountability procedure attached to the judgment, where
that process — not a single forecast — is what is under review.
Output
Per finding: name the flaw (hedgehog overconfidence mistaken for insight, unwarranted credit for a
call that a matching predisposition explains just as well, a fame- or media-driven credibility
claim, an intuitive judgment trusted outside a regular feedback-rich domain, a selection or
prediction procedure that lets impression override formula, a hindsight-tainted or win-skipping
postmortem, a blurred forecasting/leading role, an accountability design that rewards looking
decisive, an untested settleable dispute, or an unjustified rational-agent-versus-behavioral
framing choice), apply the corrective technique (test the rival explanation, discount the fame or
"almost right" claim, require prolonged-feedback evidence or a formula-combined score, redo the
postmortem to include wins and contemporaneous notes, separate the forecasting and leading voices,
propose the adversarial-collaboration test, or justify the modeling choice), state the residual
uncertainty the correction leaves, and end with a concrete next step. Order findings
highest-impact first. Never substitute a bare verdict — "reliable forecaster" or "good call" — for
this structure.
Anti-patterns to flag
- A single, confident narrative treated as insight because it "sounds sure," with no qualifying
language and no visible engagement with a rival account (P192).
- A forecaster's correct call credited to superior method or unusual openness, when the record
shows no more than a predisposition that happened to match events (P078).
- An aggregated estimate or "wisdom of crowds" claim built from inputs that are not actually
independent or diverse — echo-chamber pooling mistaken for triangulation (P073, P152).
- A track record praised without naming the baseline it beat, especially one earned by looking
confident mid-crisis rather than by beating a regime-appropriate baseline (P153); or a record
built on thin, case-by-case judgment mistaken for sufficient expertise (P154).
- A forecaster's visibility, media profile, or confident delivery cited as if it were evidence of
accuracy (P155); or open-ended "I was basically right" credit granted without discounting for
how self-serving the claim is (P156).
- A claimed expert intuition accepted at face value without checking the domain is regular and
feedback-rich (P083); a low-validity selection or forecasting procedure that lets a final human
impression override a structured score (P187); or an evaluation that rates candidates
holistically instead of trait-by-trait, inviting halo contamination (P188).
- A postmortem that skips the wins, skips contemporaneous notes, or credits a good outcome to a
good decision without ruling out luck or offsetting errors (P085); or a response to a miss that
rationalizes rather than treats the failure as data to adjust from (P121).
- A leadership decision that blurs the humble, probabilistic forecasting voice into the confident,
decisive command voice, or the reverse, instead of reconciling the two through stated intent
(P086); a decision still being re-litigated well past the point resolute execution should have
begun, or a plan followed rigidly after the situation has changed (P087).
- An accountability design that rewards looking decisive after the fact rather than sober
reasoning offered before the outcome was known (P197); a polarized dispute left untested when a
concrete, benchmarked forecast could have settled it (P122).
- A rational-agent-versus-behavioral framing choice made with no justification, or the choice
smuggled into an unrelated claim about how the actor values outcomes or trades off gains and
losses (P103).
References
See ../../references/bias-perception-principles-index.md for the full principle catalogue. For
adjacent concerns, see the sibling skills: historical-analogy-learning-and-hindsight covers the
broader hindsight-bias mechanism that this skill applies narrowly to a forecaster's own
postmortem; calibration-and-probabilistic-estimation scores the numeric probability judgment
once the track-record and process checks here are satisfied; prospect-theory-framing-and-decision-weights
owns how the reviewed actor values outcomes and trades off gains and losses — the wider territory
P103 borders here only for the rational-agent-vs-behavioral modeling choice.
Provenance
Derived solely from P073, P078, P083, P085, P086, P087, P103, P121, P122, P152, P153, P154, P155,
P156, P187, P188, P192, and P197 (Superforecasting; Expert Political Judgment: How Good Is It? How
Can We Know?; Thinking, Fast and Slow — all distillation-only; see the frontmatter above for the
full claim, evidence, and source-anchor list).