Reliability work
Incident follow-up
A calm follow-up note that turns a disruption into a better future check.
What this covers
- What users may have seen and when it was noticed.
- What changed during repair and how the result was verified.
- One prevention item that is small enough to complete.
Write for future work
An incident note should help the next person understand what happened and what changed. It does not need to become a long report or a blame exercise.
The useful version is usually compact: impact, first signal, working theory, repair action, verification, and one follow-up. That is enough for many small public services.
Keep cause and evidence separate
Early explanations are often incomplete. The note should separate observed facts from likely causes. For example, the page returned an unexpected status is an observation. A deployment change caused it is a theory until the repair confirms it.
This separation keeps the note readable and prevents future work from copying an assumption as fact.
Turn one thing into a routine check
The best prevention item is small enough to finish. If the issue was missed because nobody checked a key page after deployment, add that page to the review list. If a certificate signal was late, add a renewal check.
A follow-up note is complete when it improves the next routine check, not when it produces a long list of ideal improvements.
Follow-up checklist
- Describe visible impact in plain language.
- Record the first signal that showed something was wrong.
- List repair actions separately from assumptions.
- Record how the final state was verified.
- Choose one prevention item with an owner.