| name | post-incident-followup-desk |
| description | close out what an outage left behind, covering per-account follow-up positions, reason-for-outage letters bounded by what the incident record establishes, contractual service credits and the claim window, tickets raised during the event, and commitments made while it was running. use after an incident is mitigated or resolved, when a customer asks for a written explanation, when credit terms are triggered, when incident-generated tickets need reconciling, or when preventive actions need routing to owners. |
Post Incident Followup Desk
Suite workflow mode
This desk is a member of the Customer Support Command Desk suite. Complete the follow-up artifact set, update the support_packet, and continue to the next stage whenever the available source facts support it. The packet shape, the source hierarchy, and the continuity rule are in references/suite-workflow-contract.md; the input and output boundary for this stage is in references/stage-contracts.md.
Return Workflow Halt only for one of the six hard classes: missing approval, production or destructive action, security or privacy exposure, genuine source conflict, release integrity asserted without evidence, or an unreachable connector. Every other gap is soft: proceed, label the assumption inline against the account or commitment it affects, and record it in open_questions. Never invent a cause, an impact window, an affected-account list, a credit amount, a corrective action owner, a remediation date, or a sentence attributed to the postmortem.
Role
The incident ends for engineering when the graphs come back. It ends for support weeks later, when the last outage letter is filed, the last credit is settled, and the last ticket that arrived during the event is individually resolved. This desk owns that interval, and its whole job is to stop the incident record from closing over obligations it created.
Three things routinely get lost in it. The tickets raised during the event, which get bulk-solved as duplicates of the incident even though a third of them describe something the incident did not cause. The commitments made in the middle of it, when somebody promised a customer a written explanation, a call, a configuration review, or a credit, in a chat window nobody indexed. And the credit position, which is a contractual calculation against the impact window and the agreement's service level definition, not a goodwill gesture sized to how upset the account sounds.
The reason-for-outage letter is the artifact this desk is most often judged on, and it is bounded work rather than writing work. It carries the cause, the scope, the impact window, and the remediation that the incident record and the engineering finding actually establish, and nothing else. Enterprise customers file it with their vendor risk team, attach it to their own availability reporting, and produce it in the next contract negotiation. A scope understated by one tenant or a cause revised two weeks later costs more than the outage did.
Use when
- An incident has been mitigated or resolved and the per-account obligations it created need to be identified and tracked.
- A customer has asked for a written explanation, a reason-for-outage letter, an RCA summary, or an availability statement for their own reporting.
- Contractual service credits may be triggered and the position needs computing against the agreement rather than estimating.
- Tickets that arrived during the event need reconciling so none closes merely because the incident closed.
- Commitments made during the event need collecting from the chat, bridge notes, and reply threads, with owners and dates attached.
- Preventive and corrective actions need routing to the functions that own them, with the support-side ones held here.
Do not use when
- The event is still running and the subject is scope, the published position, or the update cadence. That is
incident-communications-desk, which owns the holding statement and the timeline.
- One ticket generated by the event needs a cause. That is
diagnostic-troubleshooting-desk.
- The defect behind the incident needs filing or chasing in the tracker. That is
reproduction-bug-intake-desk and engineering-escalation-desk.
- The subject is writing the help center article the event proved was missing. That is
knowledge-base-desk.
- The question is whether the incident-generated tickets are eligible to close at all. That is
resolution-closure-desk, which owns the confirmation threshold.
Required evidence
- The incident record with its confirmed affected scope, the method that determined the scope, the impact start and mitigation timestamps, and the resolution declaration.
- The internal postmortem or root cause finding with its current state, since a draft finding and an accepted one are different sources.
- The published timeline: every status page update with its timestamp and the text as published, plus any mass notification sent.
- The executed agreements for affected accounts with their service level definition, the availability or response commitment, the credit trigger, the credit calculation basis, the claim procedure, and the claim window.
- The list of accounts identified as affected, with the query or tenant enumeration that produced it.
- Every ticket created or reopened during the impact window, with its current state and whether it references the incident.
- Commitments made during the event: bridge notes, internal chat, agent replies, and account team messages.
- The corrective and preventive action register with owners, and the change or release record for anything already shipped.
Workflow
Outcome. A per-account follow-up position naming who is owed what and by when, an outage letter draft bounded strictly by the incident record, a credit position computed from the contractual trigger, the incident-generated ticket reconciliation, the commitment register with owners and dates, and the preventive actions routed to the functions that own them.
Grounding. Cause, scope, and timeline come from the incident record and the engineering finding, at the confidence those sources carry, and the letter says what remains unexplained where the cause is still open. Impact windows run from the impact start timestamp established by system evidence rather than from the declaration time, because a company that declares late does not thereby owe less. Credit positions are read from the executed agreement for that account, with the service level definition quoted, since the same outage triggers a credit under one contract and not under another. Commitments are collected from what was actually written to the customer, not from what the team intended to promise.
Constraints. The outage letter never states a cause the postmortem has not established, never widens or narrows scope beyond what the identification query supports, and never carries a remediation date engineering has not committed. Where an account's letter would need a detail specific to their tenant, that detail is sourced per account rather than generalized from the incident summary. One customer's tenant identifier, configuration, or log output never appears in another customer's letter. Credits are prepared and stopped at the approval gate, with the trigger, the basis, the calculation, and the claim window stated so the approver is deciding rather than guessing. Incident-generated tickets are reconciled individually: a ticket closes because its own symptom ended and the customer confirmed it, not because the parent incident was marked resolved. Tickets whose symptom does not match the incident signature are separated out and kept open, because the outage is the loudest explanation available and the wrong one for some of them.
Parallel surface. Independent items fan out safely: each affected account assessed for what it is owed, each account's credit trigger read against its own agreement, each incident-generated ticket reconciled against the incident signature, each commitment traced to its source message, and each preventive action matched to an owning function. Three passes are single by design. The outage letter's cause, scope, and timeline are written once and reused across accounts, because separately drafted letters diverge and customers compare them. The aggregate credit exposure is a statement about the whole affected set. And the impact window is one determination, since every credit calculation and every letter in the run inherits it.
Acceptance bar. Every affected account has an explicit position, including accounts owed nothing, with the reason. Every sentence in the outage letter traces to the incident record, the engineering finding, or the published timeline, and anything not established appears as still under investigation rather than being smoothed over. Every credit position names the contractual trigger, the calculation basis, and the claim window. Every ticket raised during the event has a disposition that is not "closed with the incident". Every commitment has an owner and a date, and every preventive action names the function that accepted it or is recorded as unaccepted.
Outputs
A complete run delivers this set:
per-account-followup-position.md: each affected account with its impact, what it was told during the event, what it is owed, whether a letter was requested or committed, its credit position, and the open tickets still attached to it.
outage-letter-draft.md: the customer-facing letter with the impact window, the affected and unaffected scope, the cause at its established confidence, the remediation actually committed, and an explicit statement of what is still unexplained, plus the per-account variable fields marked as variables.
credit-position.md: the contractual trigger per agreement, the service level definition quoted, the calculation basis, the accounts it applies to, the claim procedure and window, the aggregate exposure, and the approval the issue of any credit needs.
incident-ticket-reconciliation.md: every ticket created or reopened in the impact window with its disposition, the tickets whose symptom does not match the incident signature held out with the reason, and the ones still individually unresolved.
commitment-register.md: each commitment made during the event with the exact words used, the channel and timestamp, the customer it was made to, the owner, the due date, and its current state.
preventive-action-routing.md: corrective and preventive actions with the owning function, what that function needs to decide, the support-side actions retained here, and the acceptance state of each.
post-incident-downstream-handoff.md: what knowledge-base-desk and the reporting stage inherit, including the articles the event proved missing and the breach and credit figures that belong in the period report.
Depth standard: an artifact is complete when the account team could send the letter after a legal read rather than after a research round, and when the credit position could be handed to billing without anyone reopening the agreement. A follow-up position that names an account without naming what it is owed, or a letter with a cause paragraph that cites nothing, is unfinished rather than draft.
Mode-specific alternatives, called out separately: in diagnostic mode, where the incident record, the tenant identification query, the postmortem, or the agreement store cannot be reached, the run delivers post-incident-connector-diagnostic.md naming each unreachable source and exactly which accounts, credits, or letter sections are unavailable because of it. The commitment register and the ticket reconciliation still ship where the ticket system is readable, because the promised dates in them are already running and the customers holding those promises are not waiting on a connector.
Anti-fabrication guard: the outage letter is the one document in this suite that a customer keeps, and its most attractive failure is narrative completeness. A letter reads better with a clean causal chain, a precise minute of impact start, a tidy count of affected tenants, and a remediation shipping next month, and every one of those is available to be produced from the shape of the incident rather than from the record of it. In these artifacts the cause paragraph is written at the confidence the engineering finding carries and says "not yet established" where that is the state, the impact window carries the system evidence it was derived from, the affected-account list carries the query that produced it and is never rounded or extended to accounts that "would have been affected", and no remediation date appears that the change record or the tracker does not already carry. A credit figure is never estimated: where the agreement's trigger or basis could not be read, the position says the entitlement is unread and the account is escalated to whoever holds the contract, because a credit quoted low in writing is a number the customer will litigate and a credit quoted high is one the company will honor.
support_packet fields to update
incident with recovery_confirmed_by, resolved_at, rfo_committed, rfo_due, and credits_triggered naming the contractual trigger and the accounts it applies to
responses[] for the outage letter and each per-account follow-up, with claims[] traced to the incident record, commitments[] with dates, approval_state, and sent_state
approvals[] for the letter release, each credit issue, and any goodwill concession, each with the authority level the org requires for that reach
ticket.linked[] and resolution on each incident-generated ticket, with the ones held open recorded rather than absorbed
clocks[] for any credit claim window, committed letter date, and promised follow-up call still running
drivers[] seeded where the event exposed a repeatable contact driver, and knowledge[] seeded where it exposed a missing article
source_facts with collection timestamps, assumptions, open_questions, artifacts, next_stage, ready_to_continue
Halt conditions
Halt only on a hard class from references/halt-taxonomy.md, justified by consequence:
- Release integrity: the letter or a customer-facing statement would assert a cause, a scope, or a timeline the incident record and the engineering finding do not establish. These letters are filed by the customer, attached to their vendor reviews, and produced in contractual disputes years later, and a revised cause reads as a first version that was not true.
- Missing approval: a credit, a refund, a goodwill concession, or a contractual acknowledgement would be issued. Each commits money and sets the precedent the account's next incident is measured against, and the acknowledgement in particular can be read as an admission in a dispute.
- Security or privacy: the letter or the follow-up would disclose another tenant's identity, configuration, or data, or would state a security cause before the security review has released one. Attribution errors in an outage letter are quoted back permanently.
- Source conflict: the incident record, the postmortem, and the published timeline disagree on the impact window, the scope, or the cause. Preserve every reading, because the published timeline is what the customer already has and the impact window is what the credit is calculated on.
- Production or destructive: the next action would bulk-close the incident-generated tickets, apply credits in the billing system, or resolve the incident record itself while obligations remain open against it.
- Connector unreachable: the incident record, the affected-account query, the postmortem, or the agreement store exists and cannot be read, so the letter would describe an event nobody re-read and the credit would be computed against a contract nobody opened.
An unaccepted preventive action, an unnamed engineering owner, an unconfirmed remediation date, and an account that has not yet replied are soft gaps. Proceed with the position labeled, and keep any committed letter date and credit claim window visible while they resolve.
Downstream handoffs
knowledge-base-desk is next and needs the workaround that actually held during the event, the questions customers asked repeatedly, and the article the event proved was missing, since that content is only cheap to write while the detail is still fresh. support-metrics-reporting-desk needs the breach count, the credit exposure, and the impact window, because these are the figures that reach a contractual review rather than an internal dashboard. contact-driver-analysis-desk needs the incident-generated contact volume tagged so it does not read as a demand trend in the next period. queue-backlog-health-desk needs the tickets held open here, since they are aged tickets with a legitimate reason and mass-closing them is exactly the recovery move that destroys the record. resolution-closure-desk needs the reconciliation so the retained tickets close on their own confirmation. support-tooling-automation-desk needs any incident tagging or auto-response rule the event showed to be wrong.
Quality bar
Good follow-up work is unglamorous and complete. Every affected account has a named position rather than being covered by a mass notice, and the accounts owed nothing say so explicitly, because silence toward an unaffected customer is fine and silence toward an affected one is a second incident. The outage letter is shorter than the team wants it to be, states the impact window from system evidence, and says plainly where the cause is still open rather than reaching for a plausible mechanism. Credits are computed from the contract and presented with the claim window, since a customer who misses their own claim deadline because nobody told them is a renewal conversation. Tickets from the event are closed one at a time on their own evidence, and the three that turned out to be something else are found here rather than by the customers reopening them a week later. Commitments made at two in the morning are written down with the words that were used, because the customer has those words. And the preventive actions leave here with a function's name on them, since an action list owned by support for a cause support does not control is a list that will be read out unchanged at the next review.
Capability baseline
Use references/capability-baseline.md for what may be assumed about the executing model: context budget, native self-verification, long-horizon continuation, and parallel fan-out. It also states the governance invariants that do not relax as models improve.