| name | calibration-audit |
| description | Use this to compare what was expected against what actually happened, and find where confidence systematically outruns outcomes. |
Calibration Audit
The Big Idea
Calibration is a measurable behavior, not a tone of voice. Assertive
language proves nothing. Graded predictions against outcomes prove
everything.
Uncertainty Handling Inventory
When the corpus shows the subject did not know something, classify the
response: asked for evidence, researched, made assumptions, requested
confirmation, ran a small experiment, tested a hypothesis, acted
immediately, escalated to another model.
Count responses per task type. The default move under uncertainty is a
fingerprint.
Expected vs Actual
Harvest expectations implicit in plans and statements: "this should take an
hour," "this approach will work," "the bug is in X."
Grade them:
- expected difficulty vs actual difficulty
- expected solution path vs path taken
- predicted failure points vs actual failure points
- estimated time vs elapsed time, where timestamps allow
Output
Two lists, stated plainly:
- Domains where perceived certainty systematically exceeds outcomes.
- Domains where the subject undersells themselves - outcomes beat their
own expectations.
When It Backfires
- Reading confident tone as confident prediction. Separate linguistic
style from calibration explicitly.
- Grading predictions the model made. Grade only USER INITIATED
expectations.
- Demanding numeric precision the data cannot give. State uncertainty
ranges instead of inventing decimals.
One-Line Memory
Grade expectations, not grammar.