bob-postmortem
Write a blameless incident postmortem — gather timeline, root cause, impact, action items
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Write a blameless incident postmortem — gather timeline, root cause, impact, action items
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
Generate a self-contained, navigable explainer bundle for a feature of the current codebase — a sidebar of sub-concepts, detail pages, and linked animated diagrams grounded in the real code
Analyze a codebase and interactively generate an OKF (.knowledge/) bundle — asks clarifying questions, discovers packages and decisions, then writes a complete navigable knowledge catalog.
Analyzes code for unnecessary complexity, unjustified abstractions, and structural cleanup opportunities using first-principles engineering methodology
Finds bugs in existing code — nil dereferences, race conditions, resource leaks, logic errors, error handling gaps. Creates cleanup tasks for each finding.
Verifies that SPECS.md, NOTES.md, TESTS.md, BENCHMARKS.md, and documentation cross-reference cleanly against each other and against the actual code
Self-directed analyst that claims analysis tasks from a shared task list and writes findings (read-only)
| name | bob-postmortem |
| description | Write a blameless incident postmortem — gather timeline, root cause, impact, action items |
| user-invocable | true |
| category | workflow |
You are writing a blameless incident postmortem. The goal is to understand what happened, why it happened, and what concrete actions will prevent recurrence — not to assign blame.
Ask the user for the following if not provided. Ask one question at a time, not all at once:
Produce a postmortem in this structure:
# Postmortem: [Title]
**Date:** [incident date]
**Status:** [Draft / In Review / Complete]
**Severity:** [P1 / P2 / P3 / P4]
**Authors:** [names]
---
## Summary
[2–3 sentence description of what happened, impact, and how it was resolved]
## Impact
| Dimension | Detail |
| ------------------- | -------------------------------------------- |
| Duration | [start] → [resolution] ([X hours Y minutes]) |
| Detected | [detection time] — [TTD: X min] |
| Mitigated | [mitigation time] — [TTM: X min] |
| Resolved | [resolution time] — [TTR: X min] |
| Users affected | [count or percentage] |
| Services affected | [list] |
| Error budget impact | [if known] |
## Timeline
| Time (TZ) | Event |
| --------- | ---------------- |
| HH:MM | [event] |
| HH:MM | [action by whom] |
| ... | ... |
## Root Cause
[The underlying technical reason. Not the trigger — the systemic issue that allowed this to happen.]
## Trigger
[The specific event that initiated the incident]
## Detection
[How the incident was detected. Include alert name if applicable.]
## Resolution
[What was done to resolve it. Step by step if relevant.]
## What Went Well
- [thing that worked — monitoring, runbook, communication, etc.]
- ...
## What Went Wrong
- [gap or failure — missing alert, unclear runbook, toil, etc.]
- ...
## Where We Got Lucky
- [things that could have been much worse]
- ...
## Action Items
| Action | Owner | Priority | Due |
| ----------------------------- | ----------- | ---------- | ------ |
| [specific, measurable action] | [name/team] | [P1/P2/P3] | [date] |
| ... | ... | ... | ... |
## Supporting Information
[Links to dashboards, logs, alert history, Slack threads, runbooks]
Root cause vs trigger: The trigger is "the deploy at 14:32 introduced a nil pointer". The root cause is "we have no integration tests for nil config values" or "the deploy pipeline doesn't run smoke tests". Dig deeper than the trigger.
Action items must be:
Severity guide:
TTD / TTM / TTR: