How do you keep runbooks up to date?
Runbooks stay up to date when four habits are in place. Each runbook has a named owner. Runbooks get run on a schedule, not only during incidents. Every change is a versioned revision that someone reviews as a diff before it publishes. And after an incident, the people who used a runbook write back what actually worked. Without these habits, wiki runbooks drift away from the systems they describe.
Why wiki runbooks go stale
A runbook is a claim about how a system behaves, and systems change all the time. Hostnames move, flags get renamed, a service splits in two, and a cluster gets replaced. None of these changes touch the wiki page, so nothing tells anyone the page is now wrong.
Several things make it worse in most organizations:
- Nobody owns the page. The author has changed teams, so the page belongs to everyone and nobody maintains it.
- The runbook only gets read during incidents. That’s the worst time to find a broken step, and the person who finds it is too busy to fix the page.
- Edits are invisible. Wiki history exists, but nobody reviews it. A quick edit at 3 a.m. becomes the official procedure with no second reader.
- Commands get copied and adapted. Engineers paste a command into a terminal, change it to make it work, and never update the source. The working version stays in someone’s shell history.
- There’s no record of what ran. If you can’t see that a runbook was used last Tuesday and step 4 failed, you can’t tell that it needs fixing.
So staleness isn’t really a discipline problem. It happens because the page has no feedback loop to the system it describes. Each practice below adds part of that loop.
Give every runbook a named owner
Assign each runbook to one team, and one person within it, accountable for its accuracy. Usually that’s the team that owns the service. Ownership should cover three things:
- Keeping commands and steps correct when the service changes.
- Reviewing proposed changes from other teams.
- Retiring the runbook when the procedure no longer applies.
Put the owner and the date of the last verified run at the top of the runbook, where responders will see them. When a team is reorganized, move its runbook ownership along with its services. A useful audit: list every runbook whose owner has left the company or whose owning team no longer exists. That list is usually where your stalest pages are.
Set a review cadence tied to risk
A single quarterly review for every page usually gets skimmed. Base the cadence on how risky the procedure is and how often its system changes:
- High risk (failovers, data restores, credential rotation, anything destructive): review and run on a short fixed cycle, and again after any architecture change to that system.
- Medium risk (scaling, restarts, routine maintenance): review on a longer cycle, or whenever the owning service ships a significant change.
- Low risk (read-only diagnostics, lookups): review when someone reports a problem.
Also trigger reviews from events, not only the calendar. A migration, a renamed service, or a deprecated tool should open a review task for every runbook that references it. Searching your runbooks for hostnames, cluster names and CLI tools is crude, but it catches most of them.
Run them regularly
Reading a runbook will never tell you as much as running it. A runbook nobody has run in six months is a guess. Ways to keep them exercised:
- Game days and drills. Pick a runbook, run it in staging or against a controlled fault, and write down every place where reality differed from the page.
- Run diagnostic steps for real. Read-only steps such as status checks, queries and log searches are cheap to run on a schedule and fail loudly when something has moved.
- Make new on-call engineers run them. Onboarding is a natural time to work through the most important runbooks with fresh eyes. New engineers notice the steps that only make sense if you already know the answer.
- Track the last successful run. If runbooks are executed in a tool that records each run, “last run” and “where it failed” become data you can see, not something people have to remember.
The trade-off is cost. Drills for destructive procedures take time and need a safe environment. Spend that effort on the runbooks you would most regret finding broken.
Version every change and publish revisions
Treat a runbook like code. Each meaningful change produces a new revision, and the version people actually run is a specific published revision, not whatever the page says right now. This gets you three things:
- You can tell exactly which version someone followed during an incident.
- You can roll back a bad edit.
- An approval for a risky step can point to a known, unchanging set of instructions.
Some teams keep runbooks as Markdown in a Git repository to get this. That works if your responders are comfortable with pull requests and you accept that the runbook and its execution live in different places. Whatever you use, make sure a published revision can’t be edited in place. If it can, “version 12” means nothing.
Review only what changed
Reviews get skipped when reviewers have to reread a whole runbook to find a one-line change. Show reviewers a diff: the steps added, removed or modified, with the old and new commands side by side. That makes review quick enough to actually happen.
What to look for in a runbook diff:
- Changed commands. Check flags, targets, and whether a read-only command became a write.
- Removed safety steps. Look for a deleted confirmation, backup or pre-check.
- New variables or hardcoded values. An environment name or account ID that should be a parameter.
- Ordering changes. Moving a step before or after a dependency it relies on.
Require a reviewer other than the author for anything that touches production state. This is the same rule as code review, for the same reason.
Capture what worked after incidents
Incidents are when runbooks are most accurate in practice and least accurate on the page. The responder has just found out which steps were wrong and which commands actually fixed the problem. That knowledge disappears unless you capture it on purpose.
Make it part of the incident process:
- Keep the commands. Save the terminal session or shell history from the response, not a summary written from memory the next day.
- Draft the change while it’s fresh. The responder edits the runbook within a day of the incident, or writes a new one if none existed.
- Separate what worked from what was tried. Dead ends belong in the incident record. Only the working path goes into the runbook.
- Route it through normal review. An incident-driven change is still a change. The owner reviews the diff before it publishes.
- Add it as a postmortem action item. “Update runbook X” should be tracked the same way as any other follow-up.
The hardest step is the first one. Rebuilding commands from memory is where errors creep back in, so capture the raw session whenever you can.
What to look for in tooling
Most of these practices can be done by hand on top of a wiki, but each manual step is one people skip. If you’re evaluating tools, check whether they:
- Record who ran which revision, and what happened at each step.
- Make published revisions unchangeable, and require review before publishing.
- Show reviewers diffs, not whole documents.
- Turn terminal sessions or shell history into draft steps.
- Allow per-run approval for risky steps, by someone other than the person running them.
Runspace, a Gravityloop product currently in private pilots, is built around this loop. Runbook steps run inline, with output streaming underneath each step. Changes are reviewed as diffs and approved before they publish. Published revisions can’t be changed. And terminal capture turns an incident’s session into a runbook draft for the team to review.
A starting checklist
If your runbooks live in a wiki today, start small:
- List every runbook and assign an owner. Flag pages without one.
- Rank them by risk and pick the top ten.
- Run each of those ten in a safe environment this quarter and fix what breaks.
- Require a second reviewer for any change to a high-risk runbook.
- Add “update or create the runbook” to every postmortem template.
These five steps won’t fix every page, but they give your most important runbooks an owner, a recent verified run, and a way to improve after each incident.
FAQ
How often should runbooks be reviewed?
Base it on risk. Review destructive or high-impact procedures such as failovers and restores on a short fixed cycle and after any architecture change. Review routine procedures on a longer cycle or when the service changes significantly. Read-only diagnostics can wait until someone reports a problem.
Who should own a runbook?
Usually the team that owns the service the runbook operates on, with one named person accountable for keeping it accurate, reviewing changes, and retiring it when it no longer applies. Move ownership along with services during reorganizations.
Is keeping runbooks in Git enough to keep them current?
Git gives you versioning and diff review, which helps. But it doesn't run the runbook or record what happened when someone used it. You still need regular runs and a way to bring incident learnings back into the runbook.
How do you update runbooks after an incident?
Save the actual terminal session or commands from the response. Have the responder draft the change within a day, keeping only the path that worked. Then route it through normal review and track it as a postmortem action item.