用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/a5c-ai/babysitter --skill calibration-trainer命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Reference for querying the Atlas knowledge graph through its MCP tools — the SECONDARY enrichment/comparison layer that adds best-practice context to systems you have ALREADY scanned from your real sources (`az`, repos, dirs). Use when you need to look up nodes, edges, kinds, clusters, stats, or wiki pages in Atlas to compare against your real inventory. (atlas graph, query atlas, atlas mcp, search the graph, graph neighbors, atlas record, atlas kinds, enrichment layer)
Atlas turns your STATED NEED into a real systems atlas by SCANNING your actual sources (Azure via `az`, git repos, local dirs) and process/data mining them, THEN enriching against the Atlas knowledge graph. Use this skill when asked to inventory/map your real systems, scan your cloud + repos + directories, mine the real processes or data they contain, or collect their real constraints/gotchas. (atlas, scan my systems, inventory our azure account, map my repos, real systems atlas, process mining, data mining, collect nuances, system discovery)
This skill should be used when the user asks to "find skills in the wild", "assimilate popular workflows", "discover SKILL.md files in repos", "research external skills", "find workflow patterns", "survey the skill landscape", "what skills exist out there", or wants to investigate public repositories for extractable processes, babysitter plugins, and reusable procedural insights. Searches GitHub for SKILL.md files, classifies repos by archetype, and maintains structured research under docs/reference-repos/.
正在显示 SKILL.md
基于 SOC 职业分类
| name | calibration-trainer |
| description | Probability calibration training skill for improving forecast accuracy and reducing overconfidence |
| allowed-tools | ["Read","Write","Glob","Grep","Bash"] |
| metadata | {"specialization":"decision-intelligence","domain":"business","category":"collaboration","priority":"medium","tools-libraries":["numpy","matplotlib","custom quiz engines"]} |
| graph | {"domains":["domain:business-intelligence"],"skillAreas":["skill-area:statistical-analysis","skill-area:data-analysis","skill-area:quantitative-modeling"],"roles":["role:data-scientist","role:data-analyst","role:research-scientist"]} |
The Calibration Trainer skill provides capabilities for assessing and improving forecaster calibration. It helps decision-makers align their confidence levels with actual accuracy, reducing overconfidence and improving the quality of probabilistic judgments.
# Generate calibration quiz
quiz_config = {
"type": "general_knowledge",
"format": "confidence_interval",
"questions": 20,
"confidence_levels": [50, 80, 90], # percentiles to elicit
"difficulty": "medium",
"domains": ["business", "economics", "technology", "geography"]
}
# Example question
quiz_question = {
"id": "Q001",
"question": "In what year was Amazon founded?",
"actual_answer": 1994,
"format": "numeric_interval",
"required_responses": [
{"confidence": 50, "prompt": "Give your best estimate"},
{"confidence": 80, "prompt": "Give a range you're 80% confident contains the answer"},
{"confidence": 90, "prompt": "Give a range you're 90% confident contains the answer"}
]
}
# Collect responses
responses = {
"participant": "John Smith",
"date": "2024-01-15",
"questions": [
{
"question_id": "Q001",
"responses": {
"point_estimate": 1997,
"interval_80": [1995, 2000],
"interval_90": [1992, 2002]
}
}
# ... more questions
]
}
# Analyze calibration
calibration_analysis = {
"participant": "John Smith",
"n_questions": 20,
"by_confidence_level": {
"80%_intervals": {
"expected_hit_rate": 0.80,
"actual_hit_rate": 0.55,
"calibration_gap": -0.25,
"interpretation": "overconfident"
},
"90%_intervals": {
"expected_hit_rate": 0.90,
"actual_hit_rate": 0.70,
"calibration_gap": -0.20,
"interpretation": "overconfident"
}
},
"brier_score": 0.18, # lower is better, 0 = perfect
"overconfidence_index": 0.23,
"recommendations": [
"Widen confidence intervals by ~25%",
"Practice with domain-specific questions",
"Use reference class thinking"
]
}
# Calibration training program
training_program = {
"participant": "John Smith",
"baseline_calibration": 0.55, # hit rate for 80% intervals
"target_calibration": 0.75,
"exercises": [
{
"week": 1,
"focus": "interval_widening",
"exercise": "Practice giving intervals 50% wider than instinct",
"quiz_count": 10
},
{
"week": 2,
"focus": "reference_class",
"exercise": "For each estimate, identify a reference class first",
"quiz_count": 10
},
{
"week": 3,
"focus": "decomposition",
"exercise": "Break complex estimates into components",
"quiz_count": 10
},
{
"week": 4,
"focus": "consolidation",
"exercise": "Apply all techniques, track improvement",
"quiz_count": 20
}
]
}
# Track progress over time
progress_data = {
"participant": "John Smith",
"history": [
{"date": "2024-01-01", "hit_rate_80": 0.55, "brier_score": 0.22},
{"date": "2024-01-15", "hit_rate_80": 0.62, "brier_score": 0.19},
{"date": "2024-02-01", "hit_rate_80": 0.68, "brier_score": 0.16},
{"date": "2024-02-15", "hit_rate_80": 0.74, "brier_score": 0.13}
],
"trend": "improving",
"improvement_rate": "4% per session"
}
{
"operation": "quiz|analyze|train|track",
"quiz_config": {
"type": "string",
"format": "string",
"questions": "number",
"confidence_levels": ["number"]
},
"responses": {
"participant": "string",
"questions": ["object"]
},
"training_config": {
"target_calibration": "number",
"duration_weeks": "number"
}
}
{
"quiz": {
"questions": ["object"],
"total_count": "number"
},
"calibration_analysis": {
"by_confidence_level": "object",
"brier_score": "number",
"overconfidence_index": "number",
"calibration_curve": "object"
},
"recommendations": ["string"],
"progress": {
"history": ["object"],
"trend": "string",
"target_achieved": "boolean"
}
| Metric | Formula | Interpretation |
|---|---|---|
| Hit Rate | % of intervals containing true value | Should match confidence level |
| Brier Score | Mean squared error of probabilities | Lower is better (0-1) |
| Calibration Gap | Expected - Actual hit rate | Positive = overconfident |
| Overconfidence Index | Average calibration gap | Quantifies overall bias |
A well-calibrated forecaster has:
The calibration curve plots stated confidence vs. observed accuracy.
| Technique | Description |
|---|---|
| Widen intervals | Start wider, narrow only with strong evidence |
| Reference classes | Use base rates from similar situations |
| Decomposition | Break estimates into components |
| Devil's advocate | Actively seek reasons to be less confident |
| Pre-mortem | Imagine being wrong, identify why |