← Guides

How do you turn an incident into a runbook?

To turn an incident into a runbook, record what you ran and what you learned while the incident is happening, then edit that record afterwards. Pull the commands from terminal history and the incident channel. Keep the steps that moved diagnosis or recovery forward and drop the dead ends. Add checks that tell the next responder what they should see, and have someone else review the runbook before it’s published.

Why incidents make good runbooks

Runbooks written ahead of time describe how someone thinks the system fails. Runbooks written from an incident describe how it actually failed: which dashboards lied, which command finally showed the stuck consumer group, and which restart order brought things back without a second outage.

The catch is timing. Two days after the incident, nobody remembers the exact flags they used, and the shell history on the bastion host has already rotated. Most of the work happens during the incident, not after it.

Step 1: Capture as you go

You don’t need polished notes. You need a record that’s timestamped and complete enough to rebuild from. During the incident:

  • Keep one running log. Use a dedicated incident channel, a shared doc, or a single scratch file that one person owns. Scattered DMs can’t be reconstructed later.
  • Write down findings next to the commands that produced them. “Lag on orders-consumer is 2.1M and growing” is useful. “Checked Kafka” is not.
  • Note your decisions and why you made them. “Skipped the failover because the replica was 40 minutes behind” is the most valuable line in the whole log, and it’s the one people forget to write.
  • Paste output, not paraphrases. Trim it, but keep the lines that told you something.
  • Flag the dead ends as dead ends. A short “tried X, no effect” is enough. You’ll want it in review and probably not in the runbook.

A useful habit is to give one responder the scribe role while the others work the problem. It feels like a lost pair of hands, but it pays for itself the first time the same failure happens again.

Step 2: Recover terminal and shell history

The commands people actually typed are the most reliable record you have, and also the easiest to lose. Default shell settings often drop them.

Set these up before the next incident:

  • Bash: shopt -s histappend, PROMPT_COMMAND='history -a' so each command is written immediately, and HISTTIMEFORMAT='%F %T ' so entries carry timestamps.
  • Zsh: setopt INC_APPEND_HISTORY EXTENDED_HISTORY gives you immediate writes and timestamps.
  • Whole sessions: script records a terminal session, output included, to a file. It’s crude but complete.
  • Raise history size (HISTSIZE, SAVEHIST) on bastions and jump hosts, where the important commands tend to run.

Afterwards, collect history from every host and every responder involved, then merge it by timestamp. Watch for gaps. Commands run inside kubectl exec, an SSH session to a second host, or a database console won’t show up in your local history, so ask responders directly what they ran there.

Tools can do the stitching for you. Runspace, which is currently in private pilots, can turn an incident’s terminal session or shell history into a runbook draft, so the starting point is the commands that were actually run rather than what people remember.

Step 3: Draft the steps

Start from the merged timeline and rewrite it as a procedure. Give each step four parts:

  1. Purpose: one line on what this step establishes or changes.
  2. Command: exact and copy-safe, with environment-specific values turned into variables.
  3. Expected output: what healthy looks like and what the failure looks like.
  4. Decision: what to do next depending on the result.

For example, a line from the incident log:

kubectl -n payments get pods | grep orders showed 3 of 8 in CrashLoopBackOff, logs said OOMKilled

becomes:

Check consumer pod health. Run kubectl -n $NAMESPACE get pods -l app=$CONSUMER. If any pods are in CrashLoopBackOff, run kubectl -n $NAMESPACE describe pod <pod> and look at the last state. If the reason is OOMKilled, go to step 5 (raise memory limit). Otherwise go to step 6.

Some practical rules for drafting:

  • Parameterize anything that was specific to this incident: namespaces, cluster names, hostnames, ticket IDs. If a later step needs a value from an earlier step’s output, such as a pod name or a replica lag, name it explicitly so the responder isn’t guessing.
  • Put read-only diagnostic steps first and the steps that change things later, with a clear marker between them.
  • Write down the rollback for every step that changes state, even if you didn’t need it this time.
  • Keep the order you’d want, not the order it happened in. In incidents, people often find the root cause on the fourth try. The runbook should check that first.

What to keep and what to drop

Keep Drop
Commands that confirmed or ruled out a cause Commands that only confirmed you were logged into the right host
The diagnostic that found the root cause, moved to the front Dead ends, unless they’re a common wrong first guess worth naming
Thresholds and numbers that drove a decision Raw output beyond the lines that mattered
Recovery steps with their rollback One-off workarounds tied to this incident’s exact state
Warnings about what made things worse Credentials, tokens, and customer data that landed in the log

A dead end is worth keeping when it’s the obvious first move and it’s wrong. A short note like “Restarting the consumers alone doesn’t help; the lag returns within minutes” saves the next responder 20 minutes.

Scrub the draft before anyone else sees it. Incident logs pick up secrets from pasted environment variables, connection strings, and curl commands with auth headers.

Step 4: Review and approve

Have someone who wasn’t on the incident review the runbook. Responders read the draft with the full context in their heads and skip right over gaps a newcomer would trip on. The reviewer should ask:

  • Could I run this at 3 a.m. without the incident channel open?
  • Does every step that changes state have a precondition and a rollback?
  • Are the variables obvious, and are the defaults safe?
  • Is anything here dangerous to run against the wrong environment?

The best test is a dry run of the read-only steps against staging or production. If a command fails because a label or flag has changed since the incident, it’s better to find out now.

For later edits, review the diff, not the whole document. Reviewers get more careful when they only have to look at what changed. With Runspace, runbook changes go through review before they publish, approvals show only what differs, and published revisions can’t be changed afterwards.

Step 5: Publish and keep it alive

Publish the runbook where on-call will actually find it: linked from the alert that fired, the service catalog entry, and the postmortem. A runbook nobody can find during the next page might as well not exist.

Then keep it current:

  • Assign an owner, a team rather than a person.
  • Link it from the postmortem’s action items so it’s tracked like any other fix.
  • Update it every time it’s used. The next incident on the same service is the cheapest review you’ll get.
  • Record each time it runs and whether it worked, so stale steps show up before they fail in production.

Common mistakes

  • Writing it from memory a week later. By then the exact commands are gone, and the runbook says “check the queue” instead of telling you how.
  • Publishing the timeline as the runbook. A chronological log is evidence. A runbook is a procedure, ordered by what to check first.
  • Hardcoding incident-specific values. The next failure will be in a different namespace or region.
  • Leaving out expected output. Without it, responders can’t tell whether a step worked.
  • Skipping review because the author was there. The author was there, and that’s exactly why they can’t see the gaps.
  • Letting it rot in a wiki. Copy-paste runbooks drift silently. If nobody records when a runbook runs and what happened, nobody notices it’s broken until it’s needed.

FAQ

When should you write a runbook after an incident?

Draft it within a day or two, while shell history and memories are fresh, and ideally before the postmortem review so the runbook can be checked alongside it. Capturing commands and findings during the incident makes this much faster.

Should dead ends from the incident go into the runbook?

Usually not. Keep a dead end only when it's the obvious first move and it's wrong, and then add a short warning so the next responder doesn't repeat it. Keep the full list in the postmortem.

How do you make sure shell history captures incident commands?

Configure shells to append history immediately with timestamps (histappend and PROMPT_COMMAND='history -a' in Bash, INC_APPEND_HISTORY and EXTENDED_HISTORY in Zsh), raise history size on bastion hosts, and use script to record whole sessions when needed.

Who should review an incident runbook?

Someone who wasn't on the incident. They'll spot missing context, unclear variables, and missing rollbacks that responders skip over because they already know the background.