Kubernetes deployment rollback runbook
This runbook rolls a Kubernetes Deployment back to its last healthy revision with kubectl. Use it when a new rollout is stuck, pods are crash-looping or failing readiness, or error rates rose right after a deploy. It starts with read-only checks, makes one approved change, then verifies the result or escalates.
When to use this
kubectl rollout statushangs or reportsexceeded its progress deadline.- New pods show
CrashLoopBackOff,ImagePullBackOff, or never become Ready. - Alerts fired within minutes of a deploy and the previous version was healthy.
Do not use it when a GitOps controller (for example Argo CD or Flux) manages the Deployment. The controller will reapply the bad spec. Revert the change in Git instead. This runbook also does not cover database migrations, or ConfigMap and Secret changes. rollout undo reverts only the pod template.
Before you start
- Access: a kubeconfig context for the target cluster with
get,listandwatchon deployments, replicasets, pods and events, pluspatchon deployments in the namespace. - Tools:
kubectlwithin one minor version of the cluster (kubectl version). - Communication: an incident channel and the owning team’s on-call contact.
Variables
{{context}}: kubeconfig context for the affected cluster.{{namespace}}: namespace of the Deployment.{{deployment}}: Deployment name.{{selector}}: label selector for its pods, such asapp=checkout.{{target_revision}}: the healthy revision to restore, taken from Step 3.
Steps
1. Confirm cluster and context
Purpose: make sure every later command hits the right cluster.
kubectl config use-context {{context}}
kubectl config current-context
kubectl -n {{namespace}} get deployment {{deployment}}
Expected: the current context prints {{context}} and the Deployment is listed. Decision: if either is wrong, stop and fix access before going on.
2. Check rollout health
Purpose: confirm the rollout is actually bad, and see how.
kubectl -n {{namespace}} rollout status deployment/{{deployment}} --timeout=60s
kubectl -n {{namespace}} get rs -l {{selector}} -o wide
kubectl -n {{namespace}} get pods -l {{selector}} -o wide
kubectl -n {{namespace}} get events --sort-by=.lastTimestamp | tail -n 30
Expected: a healthy rollout prints successfully rolled out. A bad one times out, the newest ReplicaSet has fewer ready pods than desired, and the events show back-off, failed probes or image pull errors. Decision: if the rollout is healthy and the symptoms are elsewhere, stop and escalate. Otherwise continue.
3. Inspect the failing pods and rollout history
Purpose: capture evidence and choose {{target_revision}}.
POD=$(kubectl -n {{namespace}} get pods -l {{selector}} --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
kubectl -n {{namespace}} describe pod "$POD" | tail -n 40
kubectl -n {{namespace}} logs "$POD" --all-containers --previous --tail=100
kubectl -n {{namespace}} rollout history deployment/{{deployment}}
kubectl -n {{namespace}} rollout history deployment/{{deployment}} --revision={{target_revision}}
Expected: rollout history lists revisions, highest is current. CHANGE-CAUSE is often <none>, so check the pod template for each candidate revision, especially the image tag. --previous errors if the container hasn’t restarted yet, which is fine to ignore. Decision: set {{target_revision}} to the most recent revision you know was healthy, usually current minus one.
4. Preview the rollback
Purpose: see the exact spec that would be applied, without changing anything.
kubectl -n {{namespace}} rollout undo deployment/{{deployment}} --to-revision={{target_revision}} --dry-run=server -o yaml | grep -E 'image:|replicas:'
kubectl -n {{namespace}} get deployment {{deployment}} -o jsonpath='{.spec.paused}{"\n"}'
Expected: the images match the healthy version. The paused field prints empty or false. Decision: if the Deployment is paused, rollout undo will refuse. Include kubectl rollout resume in the approval request for Step 5 rather than running it now.
5. Roll back
Purpose: restore the healthy pod template.
Requires approval: this replaces running production pods with the pod template from {{target_revision}}.
kubectl -n {{namespace}} rollout undo deployment/{{deployment}} --to-revision={{target_revision}}
kubectl -n {{namespace}} rollout status deployment/{{deployment}} --timeout=300s
Expected: deployment.apps/{{deployment}} rolled back, then successfully rolled out. The restored template gets a new, higher revision number. Decision: if status times out, go to “Roll back or escalate”.
Verify
kubectl -n {{namespace}} get deployment {{deployment}} -o jsonpath='{.spec.template.spec.containers[*].image}{"\n"}'
kubectl -n {{namespace}} get deployment {{deployment}}
kubectl -n {{namespace}} get pods -l {{selector}}
kubectl -n {{namespace}} rollout history deployment/{{deployment}}
Check four things. The images match the healthy version. READY equals desired replicas. Pods stay Running with no new restarts for at least five minutes. Error-rate and latency dashboards return to their pre-deploy baseline. Then post the restored revision and image in the incident channel and block the bad version from redeploying until it’s fixed.
Roll back or escalate
- Rollback stalls or new pods also fail: the cause is probably not the image. Suspects include a changed ConfigMap or Secret, a dependency, quotas or node capacity. Don’t keep cycling revisions. Escalate to the owning team with the output from Steps 2 and 3.
- Pods are healthy but errors continue: the problem is outside this Deployment. Escalate to the incident commander.
- Rolled back to the wrong revision: repeat Step 5 with the correct
{{target_revision}}. Read the revision numbers fromrollout historyagain first, because the rollback created a new one.
Any escalation should include the context, namespace, the bad and restored revisions, the images, pod events and logs, and the times you ran each step.
Running this runbook in Runspace
In Runspace each step runs inline, and its output streams under the step, so nobody pastes commands from a wiki page. {{target_revision}} can be captured from Step 3’s output. Step 5 runs only after a separate reviewer approves it, and that approval is pinned to the exact command and runbook revision. Each run is recorded in a server-side audit log. Runspace is in private pilots. Request a pilot with your work email and we’ll reach out to set it up with your team.
FAQ
Does kubectl rollout undo revert ConfigMaps and Secrets?
No. It restores only the Deployment's pod template from a previous ReplicaSet. If the bad change was in a ConfigMap, Secret or another resource, revert that resource separately.
Why does the revision number go up after a rollback?
Kubernetes doesn't restore the old revision in place. It copies that revision's template into the current spec and records the result as a new, higher revision. Run rollout history again before you pick another target.
Can I roll back a paused Deployment?
No. kubectl refuses to roll back a paused Deployment. Resume it with kubectl rollout resume first, and treat the resume as part of the approved change.
Should I use kubectl rollout undo if Argo CD or Flux manages the Deployment?
Generally no. The GitOps controller will sync the Deployment back to whatever is in Git. Revert the commit, or roll back through the controller.