What is an executable runbook?
An executable runbook is an operational procedure where each step’s command runs directly from the document, and its output appears right under that step. Instead of copying commands from a wiki page into a terminal, an operator reads the instructions, fills in variables, runs each step in place, and leaves behind a record of what ran, with which values, and what came back.
Executable runbooks vs. wiki runbooks
Most production runbooks today are static pages: a Confluence or Notion doc, a Markdown file in a repo, sometimes a Google Doc linked from an alert. They describe what to do, but nothing connects the page to the work. The operator copies a command, pastes it into a terminal, edits a hostname by hand, reads the output in another window, and decides what to do next.
That gap causes predictable problems:
- Drift. The page says
--region us-east-1; the service moved two quarters ago. Nobody notices until someone follows the page during an incident. - Copy-paste errors. A placeholder like
<CLUSTER>gets pasted literally, or the wrong environment’s value is substituted. - No record. After the incident, nobody can say exactly which commands ran, in what order, against what, or what the output was. Reconstructing it means scrolling through someone’s shell history.
An executable runbook keeps the same human-readable structure (context, warnings, checklists, decision points) but makes each command a runnable step. The document and the execution are the same artifact.
| Wiki / static runbook | Executable runbook | |
|---|---|---|
| Commands | Copied into a terminal | Run in place, step by step |
| Output | Lives in someone’s terminal | Streams under the step that produced it |
| Inputs | Hand-edited placeholders | Declared variables, filled once |
| Record of a run | Usually none | Each run is recorded |
| Freshness | Breaks silently | Breaks visibly when a step fails |
How inline execution works
The model is close to a notebook. A runbook is a sequence of steps. Each step has prose explaining what it does and why, and, where relevant, a command. The operator runs a step, and its output streams underneath as it executes, the same way it would in a terminal. Long-running commands show progress live rather than appearing all at once at the end.
Because output sits next to the step that produced it, the operator can check the result against what the runbook says to expect (“you should see three replicas in Ready”) before moving on. Steps that are purely manual, such as “confirm in the status channel that customer support has been notified,” stay as checklist items alongside the commands.
Variables, including values from earlier steps
Variables replace hand-edited placeholders. A runbook declares its inputs, such as namespace, cluster or incident_id, and the operator sets them once at the start of a run. Every step that references them gets the same value, which removes a common source of wrong-environment mistakes.
The more useful pattern is capturing a value from an earlier step’s output. For example, in a database failover runbook:
- Find the current primary.
kubectl get pods -n {{namespace}} -l role=primary -o nameThe output is captured asprimary_pod. - Check replication state on the primary.
kubectl exec -n {{namespace}} {{primary_pod}} -- pg_controldata /var/lib/postgresql/data - Cordon the node running it. The node name is captured from step 2’s context and used in the next command.
Nobody retypes a pod name from one window into another. The runbook carries the value forward, and the recorded run shows exactly which pod each command targeted.
Why teams adopt them
Teams usually move to executable runbooks for one of three reasons.
Incidents go faster and more safely. Fewer manual steps between reading and doing means fewer transcription errors under pressure, and less time spent finding the right page, the right terminal and the right credentials.
Runbooks stay accurate. When a runbook is executed, a broken step fails visibly, often during a routine run rather than the next outage. Staleness becomes something you notice and fix, not something you discover at 3 a.m.
There is finally a record. Post-incident reviews, change management and compliance all ask the same question: what was actually done? An executable runbook answers it with the commands, the values, the output and who ran them, rather than with recollection.
There is also a knowledge-transfer benefit. A newer engineer can run a procedure they have never seen before with the reasoning on the page and the results in front of them, instead of relying on the one person who knows the commands by heart.
What to look for in an executable runbook tool
Running commands from a document is the easy part; several open-source tools do it well. For a team running production at scale, the harder questions are about control and accountability. These are the capabilities worth checking.
Approval on risky steps. Not every step needs sign-off, but a failover, a data migration or a bulk delete should. Look for per-run approval by a separate reviewer, not self-approval. The approval should be pinned to the exact command and the exact runbook revision being run. If the command or the runbook changes after approval, the approval should no longer apply.
Runbook review before publishing. Changes to a runbook are changes to how production gets operated, and they deserve review the way code does. Look for a workflow where edits are reviewed and approved before they publish, and where the reviewer sees only what changed rather than rereading the whole document.
Immutable published revisions. Once a revision is published, it should not be editable in place. That is what makes “approved for revision 14” and “ran revision 14” mean something.
A server-side audit log. The record of who ran what, where, and who approved it should live on the server, not in a browser tab or a local file that can be lost or edited.
Role-based access. Teams share runbooks, but not everyone should be able to edit, run or approve every one. Roles and restricted runbooks let you open up routine procedures while keeping sensitive ones limited.
A way to get runbooks written in the first place. The biggest obstacle is often authoring. Tools that turn terminal history or an incident’s terminal session into a runbook draft let you capture what worked while it is fresh, then send it through review.
Runspace, from Gravityloop, is built around this governance layer: inline execution with streamed output and variables, plus per-run approval pinned to command and revision, reviewed changes, unchangeable published revisions, a server-side audit log, and terminal capture for drafting runbooks from incidents.
Trade-offs to consider
Executable runbooks are not free to adopt.
- Execution context matters. Commands need to run somewhere with the right network access and credentials. Decide early where steps execute and how that is controlled.
- Not every procedure should be automated end to end. Many runbooks are mostly judgment calls with a few commands. Keep the prose and decision points; making steps runnable should not turn a runbook into a script with no explanation.
- Approval adds latency. Requiring a second person on every step slows things down during an incident. Reserve approval for steps where a mistake is costly, and make requesting it fast.
- Migration takes effort. Converting a wiki full of runbooks is real work. Start with the procedures that run most often or carry the most risk, not the whole library.
Getting started
A practical path for most platform and SRE teams:
- Pick three to five runbooks that are run often or that have caused trouble in past incidents.
- Convert their commands into runnable steps, and replace placeholders with declared variables.
- Mark the steps that should require approval, and decide who can approve them.
- Run them in non-production first, fix what breaks, and publish through review.
- After the next incident, turn the terminal session into a draft runbook instead of writing it up from memory.
Runspace is currently in private pilots with a small number of teams. If you want to try this approach with your team’s runbooks, you can request a pilot with your work email, and we’ll reach out to set it up.
FAQ
What is the difference between an executable runbook and a script?
A script runs start to finish with no human in the loop. An executable runbook keeps the explanation, checklists and decision points of a written procedure, and lets an operator run each step individually, check its output, and decide what to do next.
Do executable runbooks replace incident response tooling?
No. They hold the procedures your team follows during and outside incidents. Alerting, paging and status communication remain separate concerns.
What should require approval in an executable runbook?
Steps where a mistake is costly or hard to reverse, such as failovers, data migrations, deletes and production config changes. Approval should come from a separate reviewer and be tied to the exact command and runbook revision.
How do you keep executable runbooks from going stale?
Run them regularly, so broken steps fail visibly, and require review for changes before they publish. Capturing terminal sessions from incidents as runbook drafts also keeps procedures close to what actually works.