Use when a forward-looking judgment or contested-evidence question needs structured analytic techniques, bias checks, sourcing discipline, alternatives, and calibrated estimative language; use critical-reasoning-and-argument for general claim and warrant testing.
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Use when a forward-looking judgment or contested-evidence question needs structured analytic techniques, bias checks, sourcing discipline, alternatives, and calibrated estimative language; use critical-reasoning-and-argument for general claim and warrant testing.
Single entry skill for any judgment that ships under uncertainty. Every output of every cohort agent that contains a forward-looking claim, a contested-evidence call, or a "we conclude / we estimate / we assess" statement must pass this skill's discipline before shipping.
This skill operates aftersource-evaluation (which establishes whether a source is trustworthy) and critical-reasoning-and-argument (which tests whether the evidence supports the conclusion), and beforeexecutive-communication (which restructures the judgment for an executive reader). It is the layer that makes the engine's outputs estimative-grade, not merely sourced.
When to use
Use this skill on any output that:
Contains the words we assess / we conclude / we estimate / likely / probably / almost certainly / unlikely
Forecasts an outcome or evolution
Adjudicates between competing accounts of the same event
Rests on a single source for a load-bearing claim
Makes a recommendation under uncertainty
Will be read by a decision-maker who treats the output as actionable intelligence (board, regulator, examiner, principal investigator, partner)
If the output is purely descriptive ("the regulator filed X on date Y"), this skill is not required — source-evaluation is sufficient.
The five rules (load first)
ICD 203 nine standards. Every estimative output must satisfy all nine before shipping. Reference: references/icd-203-and-tradecraft-standards.md.
Multiple competing hypotheses. No load-bearing judgment ships from a single hypothesis without an Analysis of Competing Hypotheses (ACH) matrix run on it. Reference: references/heuer-pherson-sats.md.
Calibrated probability. Forward-looking confidence is expressed in the Kent / ODNI lexicon with explicit numeric bands. Confidence in the source and probability of the judgment are stated separately. Reference: references/kent-estimative-probability.md.
Cognitive-bias self-audit. Before shipping, the analyst names which biases the conclusion is most exposed to and what mitigation was applied. Reference: references/cognitive-bias-checklist.md.
No single-source key judgment. No key judgment may rest on a single uncorroborated stream. If only one source exists, the judgment ships as a hypothesis with a confidence label of LOW and an explicit indicator of what would corroborate it. Reference: references/sourcing-and-deception.md.
Team design / analyst development / collaboration or communication breakdown
references/behavioral-science-for-analysis.md
Pre-publication self-audit
references/cognitive-bias-checklist.md
ICD 203 — one-paragraph summary
The Office of the Director of National Intelligence's Intelligence Community Directive 203 (revalidated January 2015) requires every IC analytic product to satisfy nine standards: appropriate sourcing; explicit uncertainty; distinction between assumption and judgment; analysis of alternatives; customer relevance; logical argumentation; consistency with prior judgments (with change explained); accurate probability judgments; and proper use of visual information. The engine adopts these as a universal ship-gate for any estimative output. Source: ODNI ICD 203 PDF — https://www.dni.gov/files/documents/ICD/ICD-203.pdf. Tier 1.
The catalog formalised by Richards Heuer Jr. and Randolph Pherson (Heuer & Pherson, Structured Analytic Techniques for Intelligence Analysis, CQ Press, 2010 and later editions) gives the analyst named, runnable mini-protocols — each with a trigger, step-by-step method, output template, and known pitfalls. The engine adopts: ACH, Key Assumptions Check, Devil's Advocacy, Team A / Team B, Red Cell, What-If, High-Impact/Low-Probability, Pre-Mortem, Indicators-and-Warning, Argument Mapping. Reference: references/heuer-pherson-sats.md.
Estimative probability — one-paragraph summary
Sherman Kent (Office of National Estimates, 1951–1973) observed that vague verbal probability ("possible", "likely") was read inconsistently across readers. He proposed pairing each verbal term with an explicit numeric band. The ODNI lexicon (almost certainly ≈ ≥95%; very likely ≈ 80–90%; likely ≈ 60–80%; probably ≈ 50–70%; might ≈ 20–50%; unlikely ≈ 10–20%; very unlikely ≤5%) is a direct descendant. The engine requires every estimative claim to use the lexicon and to separate probability of the judgment from confidence in the source. Reference: references/kent-estimative-probability.md.
Universal output ship-gate
Before any estimative output ships, every box must be ticked:
ICD 203 nine standards all satisfied (run the checklist in references/icd-203-and-tradecraft-standards.md).
Critical-reasoning gate passed: argument map, countercase, inference audit, and fallacy/coherence audit completed before estimative language is applied.
At least one alternative hypothesis considered — even if rejected. ACH matrix attached for any high-stakes judgment.
Key Assumptions Check run — every load-bearing assumption listed and challenged.
Confidence vocabulary drawn from the Kent / ODNI lexicon, with numeric band attached.
Probability and source-confidence stated separately.
Cognitive-bias audit — analyst names which biases the conclusion is most exposed to and what mitigation was applied.
No single-source key judgment — or, if one exists, it is labelled HYPOTHESIS, given LOW confidence, and accompanied by indicators that would corroborate.
Change explanation — if this judgment differs from a prior judgment by the engine, the change is explained and the trigger evidence cited.
Dissents footnoted — any minority view inside the analytic team or against an external SME is recorded as a footnote, NIE-style.
If the failure mode was organizational rather than evidential, the collaboration / communication note from behavioral-science-for-analysis.md is attached.
Universal anti-patterns
Confirmation bias dressed as analysis. Cherry-picking evidence that supports the dominant hypothesis without surfacing evidence that disconfirms it.
Mirror-imaging. Assuming a counterpart (competitor, adversary, examiner) reasons the way the analyst does.
Conflating source confidence with judgment probability. "High-confidence judgment" that turns out to rest on a single uncorroborated source.
"High / medium / low" confidence with no anchor. Useful only when paired with the Kent / ODNI numeric band.
Single-source key judgment. The Iraq WMD NIE 2002 is the canonical case: the central claims rested on Curveball (HUMINT) and the Samarra trucks (IMINT), neither corroborated. The fix is structural, not editorial.
Devil's Advocate as theatre. A devil's advocate without standing or budget is a box-tick. The advocate must be empowered to surface a contrary memo that ships alongside the main judgment.
No change explanation. A judgment that contradicts a prior judgment ships without a paragraph on why. Examiners and decision-makers notice.
Word-smithing as coordination. The Bruce critique of IC coordination: it has been "corrupted into a linguistic exercise" rather than a re-test of evidence-inference linkage.
Companion skills
source-evaluation — runs first; establishes whether each source is trustworthy and at what tier.
critical-reasoning-and-argument — runs before this skill; establishes claim logic, countercase, and evidence-to-conclusion validity for any intelligence, market, risk, or uncertainty judgment.
research-orchestration — wave dispatch; this skill plugs into Wave 2 (gap-fill triggers ACH if a contested judgment emerges) and Wave 3 (verification triggers Key Assumptions Check + cognitive-bias audit).
executive-communication — runs after; restructures the estimative judgment for the senior reader. Estimative discipline survives the restructuring: the Kent lexicon and the dissent footnotes do not get smoothed away.
academic-reporting-standards — for academic outputs, the equivalent quality controls are GRADE (evidence quality), TOP (transparency), Cochrane RoB (risk of bias). The Heuer/Pherson SATs and ICD 203 are the intelligence-side analogue.
due-diligence — DD outputs are estimative judgments by another name; this skill applies in full.
consulting-delivery — for turning the judgment into a client-ready problem-solving recommendation.
Sources for this skill
Use When
Use for contested evidence, alternative hypotheses, warning, deception, or forward-looking judgments.
Do Not Use When
Use critical-reasoning-and-argument for ordinary argument review without an estimative problem.
Inputs
Input
Source/provider
If absent
Question, hypotheses, evidence, decision horizon
Analyst and verified corpus
Stop judgment and return gaps
Source reliability and confidence notes
Source evaluation
Do not substitute reliability for judgment confidence
Workflow
Frame the question and plausible alternatives.
Select the smallest fitting analytic technique.
Test evidence, assumptions, bias, deception, countercases, and indicators.
Stop on unsupported load-bearing claims; recover with a qualified gap.
Express the judgment with calibrated confidence and revision indicators.
Outputs
Artefact
Consumer
Acceptance condition
Analytic judgment and technique record
Decision-maker and reviewer
Alternatives, evidence, assumptions, confidence, indicators, and dissent are explicit
Tradecraft Evidence Guidance
The technique worksheet, evidence matrix, assumption log, and confidence rationale support the judgment.
Capability Contract
Analysis is read-only by default. Collection, surveillance, operational action, publication, or source-record changes require explicit authority.
Degraded Mode
With limited evidence or tools, provide conditional hypotheses and collection gaps; do not manufacture probability or resolution.
Decision Rules
Choice
Action
Failure/risk avoided
Multiple explanations fit
Compare hypotheses
Premature closure
Low-likelihood outcome has high impact
Add indicators and warning threshold
Risk blindness
Deception is plausible
Test source independence and motive
Manipulated judgment
Tradecraft Quality Notes
Judgments distinguish evidence, inference, assumptions, source reliability, and analytic confidence.
Tradecraft Pitfalls
Choosing a technique by habit; match the question.
Hiding alternatives; show them.
Conflating confidence with source quality; report both.
Assigning unsupported precision; qualify it.
Omitting disconfirming evidence; record it.
Tradecraft Scenario
Two explanations supported by overlapping sources remain separate until independent evidence discriminates between them.
National Research Council. Intelligence Analysis for Tomorrow: Advances from the Behavioral and Social Sciences. National Academies Press, 2011. Tier 1.
The verbatim attribution discipline applies in full: claims in the references that paraphrase George/Bruce, Heuer, Pherson, Kent are labelled as paraphrase and tied back to the canonical source. Quotations carry chapter/page where the source allows.
Evidence Produced
Evidence
Consumer
Acceptance condition
Technique worksheet, hypothesis matrix, and assumption log
Decision-maker and analytic reviewer
Each judgment separates evidence, inference, assumptions, confidence, alternatives, and indicators
Quality Standards
The judgment exposes disconfirming evidence, distinguishes source reliability from analytic confidence, and states observable indicators that would change the assessment.
Anti-Patterns
Choosing a technique by habit. Fix: match it to the estimative question.
Hiding plausible alternatives. Fix: test them explicitly.
Conflating confidence with source quality. Fix: report both dimensions.
Assigning unsupported precision. Fix: use calibrated qualified language.
Omitting disconfirming evidence. Fix: record and weigh it before judgment.
Worked Example
When two explanations depend on the same reporting chain, the analyst keeps both hypotheses open until independent evidence discriminates between them.