How do you require a second reviewer before running production commands?
To require a second reviewer before a production command runs, put an approval gate between the person who wants to run it and the system that executes it. The reviewer must be someone other than the requester, must see the exact command and target, and the approval must be recorded. Teams do this with change tickets, pull request reviews, a two-person rule, or per-run approval pinned to the command.
What a second-reviewer control has to guarantee
Whichever approach you pick, check it against five properties. If one is missing, the control is weaker than it looks.
- Separation. The approver is a different person from the requester, and the system enforces this. A policy that says “get a second pair of eyes” is not enforcement.
- Specificity. The approval covers a specific action: this command, against this target, with these parameters. Approving “restart the payments service” is not the same as approving
kubectl rollout restart deploy/payments -n prod-eu. - Binding. What gets approved is what actually runs. If someone can edit the command after approval, the review proves nothing.
- Timeliness. The approval comes shortly before execution and expires. An approval from last Tuesday should not cover tonight’s run.
- Evidence. A record survives the event: who asked, who approved, what ran, where, when, and what the output was. An engineer’s own notes don’t count. The record has to live somewhere the requester can’t edit.
Approach 1: Change management tickets
The traditional route: open a change request in your ITSM tool, describe the change, get approval from a peer or a change advisory board, then run the work during the approved window.
Strengths: Auditors know it. It works for scheduled, planned work, and it gives you a place for risk assessment and rollback plans.
Weaknesses: The ticket usually describes intent in prose, not the literal commands. Nothing stops the operator from running something different, or running the right thing against the wrong cluster. The link between “approved” and “executed” depends on trust plus whatever shell history someone pastes in afterward. During an incident, waiting for a board is not realistic, so teams fall back on emergency changes that are approved retroactively.
Approach 2: PR-based approval
Express the change as code, such as Terraform, Helm values, a migration file, or a script in a repo. Require approval from a code owner before merge, then let CI/CD apply it. Branch protection enforces separation, and the merged commit is what runs.
Strengths: Strong binding. The reviewed diff is what executes. You get history, review comments and a revision you can point to.
Weaknesses: Many operational actions don’t fit this model. Draining a node, failing over a database, rotating a credential, clearing a stuck queue or running a one-off data fix are imperative steps, often with decisions made between them. Wrapping each one in a PR-and-pipeline adds minutes you may not have, and engineers end up running “just this once” commands from a laptop outside the control. It works well for declarative infrastructure and poorly for runbooks.
Approach 3: The two-person rule
Borrowed from high-assurance environments: no single person may perform a sensitive action alone. In practice this looks like a second engineer on a call watching the screen share, a pairing requirement for production access, or split credentials where two people each hold part of the access.
Strengths: Fast to adopt, and it fits imperative work. The second person can catch mistakes in real time, not just at review.
Weaknesses: Verification is social, not technical. Nothing records what the observer actually saw or agreed to. “I was on the call” is weak evidence for an auditor, and a tired observer at 3 a.m. catches less than you’d hope. Split credentials are enforceable but clumsy, and they usually approve a session, not a command.
Approach 4: Per-run approval pinned to the command and revision
Here, the runbook step itself is the unit of approval. When an operator goes to run a step marked as risky, the system requests approval from a separate reviewer. The reviewer sees the exact command with variables resolved, the target, and the runbook revision it came from. The approval is bound to that command and that revision. If either changes, approval is required again. Execution and output are logged on the server.
Strengths: This satisfies all five properties at the granularity operations actually happen. It works for imperative, step-by-step procedures, and it’s quick enough to use mid-incident because the reviewer approves one step, not a whole change.
Weaknesses: It needs your runbooks to be executable and versioned, not wiki pages copied into a terminal. Runbook content also needs its own review, since an approved run of a bad runbook is still bad. That’s why published revisions should be immutable and changes reviewed before they publish.
This is the model Runspace is built around: approval by a separate reviewer, pinned to the exact command and runbook revision, with unchangeable published revisions and a server-side audit log. It is currently in private pilots.
Comparison
| Approach | Enforced separation | Binds approval to what runs | Fits imperative ops | Incident speed | Audit evidence |
|---|---|---|---|---|---|
| Change ticket | Usually | No | Partly | Slow | Intent, not execution |
| PR approval | Yes | Yes | Poorly | Slow to medium | Strong for declarative changes |
| Two-person rule | Social | No | Yes | Fast | Weak |
| Per-run, pinned approval | Yes | Yes | Yes | Fast | Strong |
Most organizations combine them. PRs for declarative infrastructure, change tickets for planned windows, and per-run approval for runbook steps and incident actions is a reasonable split.
Keeping the control during incidents
Incidents are where second-reviewer controls usually break, because the controls were designed for planned work. Plan for this in advance:
- Tier steps by risk. Read-only diagnostics (
kubectl get, log queries,EXPLAIN) should not need approval. Gate writes, deletes, failovers and anything touching customer data. If everything needs approval, people route around it. - Make the reviewer reachable. Approval requests should go to the on-call secondary or incident commander where they already are, and approving should take seconds, not a context switch into another tool.
- Define break-glass explicitly. Sometimes no second person is available. Break-glass access should be time-limited, alert loudly, log everything, and trigger a mandatory after-the-fact review. Track how often it’s used. Frequent break-glass means the normal path is too slow.
- Capture what actually happened. Improvised commands during an incident are the ones most likely to be wrong and least likely to be recorded. Capturing the terminal session gives you evidence for the postmortem and a starting point for turning the fix into a reviewed runbook.
What auditors look for
Frameworks such as SOC 2, ISO 27001 and PCI DSS all include change management and separation of duties expectations. The exact wording differs, but reviewers generally ask for the same evidence:
- A defined policy saying which production actions require approval and who can approve.
- System enforcement that requesters can’t approve their own actions.
- A sample of production changes, each traceable from request to approval to execution.
- Proof that the approved change matches what was executed.
- Handling of emergency changes, including retroactive review.
- Access reviews showing that approver and operator roles are limited to the right people.
- Logs stored where operators can’t modify them.
The third and fourth items are where wiki-plus-terminal workflows fail. If the only record of what ran is someone’s shell history, you can’t show that the approved action is the executed one.
Rolling it out
- Inventory risky actions. List the production commands your teams actually run, starting from recent incidents and the most-used wiki runbooks.
- Classify them. Mark each one as read-only, reversible write, or irreversible/high-impact. Require approval for the last two.
- Pick a mechanism per class. Use PRs for declarative changes. For imperative steps, use executable runbooks with per-run approval.
- Define approvers with roles. Decide who can run, who can approve, and which runbooks are restricted. Enforce this through your identity provider, not a shared list.
- Write the break-glass procedure and test it before you need it.
- Review the evidence monthly. Pull a sample of runs and check that request, approval and execution line up. Fix gaps before an auditor finds them.
FAQ
Can a pull request review count as a second reviewer for production commands?
Yes, for changes expressed as code and applied by a pipeline, because the reviewed diff is what runs. It fits imperative operational steps such as failovers, node drains or one-off data fixes poorly, and those often end up run outside the control.
Is a teammate watching a screen share enough to meet a two-person rule?
It catches some mistakes, but nothing enforces it and it leaves weak evidence. Nothing records what the observer saw or approved, so it rarely satisfies an auditor asking to trace a specific production change from approval to execution.
How do you require approval during an incident without slowing response?
Gate only risky steps, not read-only diagnostics. Route approval requests to the on-call secondary or incident commander. Keep approvals to a single step. Define a time-limited, logged break-glass path with mandatory review afterward.
What does per-run approval pinned to the command mean?
Each execution of a risky step needs approval from a separate reviewer. The approval is tied to the exact command and the runbook revision it came from, so if either changes, the run has to be approved again.