High 5xx error rate runbook
This runbook resolves a sustained spike in HTTP 5xx responses from a service running on Kubernetes. Use it when an error-rate alert fires or users report failures. The read-only steps come first: confirm impact, then check health, recent deploys, logs and dependencies. After that you mitigate by rolling back or scaling, and verify that the service has recovered.
When to use this
- The 5xx ratio for one service has been above your alert threshold for at least 5 minutes.
- Users report errors and synthetic checks fail on that service.
- A deploy went out and errors rose right after.
If several unrelated services are failing at once, suspect shared infrastructure (ingress, DNS, the cluster itself) and escalate before you work through this runbook.
Before you start
kubectlaccess to the namespace, with rights to read pods, logs, events and deployments. Steps 6 and 7 also need rights to patch deployments and HPAs.curlandjqinstalled locally.- Network access to the Prometheus HTTP API.
- A read-only Postgres role. Supply credentials through
PGPASSWORDor~/.pgpass, never inline. - The incident channel open, so you can post what you find as you go.
Variables
{{namespace}}: Kubernetes namespace of the service.{{deployment}}: deployment name, e.g.checkout-api.{{selector}}: pod label selector, e.g.app=checkout-api.{{service}}: the service label your metrics use.{{prometheus_url}}: Prometheus base URL, e.g.http://prometheus.monitoring:9090.{{service_url}}: base URL of the service’s health endpoint.{{upstream_health_url}}: health URL of the upstream dependency you suspect.{{db_host}},{{db_name}},{{db_user}}: Postgres connection details for the read-only role.{{good_revision}}: last known-good rollout revision (you find it in step 3).{{replicas}}: target replica count for scaling.
Steps
1. Confirm impact
Purpose: measure the actual 5xx ratio before you act on it.
curl -sG "{{prometheus_url}}/api/v1/query" \
--data-urlencode 'query=sum(rate(http_requests_total{service="{{service}}",code=~"5.."}[5m])) / sum(rate(http_requests_total{service="{{service}}"}[5m]))' \
| jq -r '.data.result[0].value[1]'
Expect a decimal between 0 and 1 (0.12 means 12% of requests are failing). The metric and label names here follow common Prometheus client conventions, so change them to match your own instrumentation. Decision: if the ratio is below threshold and still falling, keep watching instead of acting. If it is above threshold, continue.
2. Check health endpoints and pod state
Purpose: find out whether pods are up, ready and staying up.
curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' "{{service_url}}/healthz"
kubectl get pods -n {{namespace}} -l {{selector}} -o wide
kubectl get pods -n {{namespace}} -l {{selector}} \
-o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.containerStatuses[0].restartCount}{"\t"}{.status.containerStatuses[0].lastState.terminated.reason}{"\n"}{end}'
Expect 200 and every pod Running and READY. Decision: OOMKilled or rising restart counts point to resource limits or a memory leak. If only some pods are unready, check which nodes they run on.
3. Check recent deploys
Purpose: line up when the errors started against when the last change rolled out.
kubectl rollout history deployment/{{deployment}} -n {{namespace}}
kubectl get rs -n {{namespace}} -l {{selector}} --sort-by=.metadata.creationTimestamp \
-o custom-columns=NAME:.metadata.name,CREATED:.metadata.creationTimestamp,REPLICAS:.status.replicas,IMAGE:.spec.template.spec.containers[0].image
Expect a list of revisions and the ReplicaSet created by each one. Decision: if the newest ReplicaSet was created shortly before errors began, write down the revision before it as {{good_revision}} and treat rollback (step 6) as the likely fix.
4. Read logs and events
Purpose: find the error the service itself is reporting.
kubectl logs -n {{namespace}} -l {{selector}} --since=15m --tail=500 --prefix \
| grep -Ei 'error|exception|timeout|refused|5[0-9]{2}' | tail -50
kubectl get events -n {{namespace}} --sort-by=.lastTimestamp | tail -20
Expect stack traces or repeated messages. Decision: connection pool exhaustion or connection refused sends you to step 5. Application exceptions that started at deploy time send you to step 6. Probe failures or FailedScheduling events point to capacity, so go to step 7.
5. Check dependencies
Purpose: rule the database pool and upstream services in or out.
psql -h {{db_host}} -U {{db_user}} -d {{db_name}} -c \
"SELECT state, count(*) FROM pg_stat_activity WHERE datname = '{{db_name}}' GROUP BY state;"
psql -h {{db_host}} -U {{db_user}} -d {{db_name}} -c "SHOW max_connections;"
curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' "{{upstream_health_url}}"
Expect total connections well below max_connections, plus a fast 200 from the upstream. Decision: if connections are near the limit or there are many idle in transaction sessions, the database side is saturated. Scaling the app (step 7) will make that worse, so escalate to the database owner. If the upstream is slow or failing, escalate to its owner and think about shedding load.
6. Mitigate: roll back
Purpose: put the last known-good revision back into service.
Requires approval: changes the production workload. Have a second engineer confirm {{good_revision}} first.
kubectl rollout history deployment/{{deployment}} -n {{namespace}} --revision={{good_revision}}
kubectl rollout undo deployment/{{deployment}} -n {{namespace}} --to-revision={{good_revision}}
kubectl rollout status deployment/{{deployment}} -n {{namespace}} --timeout=5m
Expect successfully rolled out. Decision: go to Verify.
7. Mitigate: scale out
Purpose: add capacity when errors come from load rather than a bad change.
Requires approval: changes production capacity and cost, and raises connection load on dependencies.
kubectl get hpa -n {{namespace}}
kubectl patch hpa {{deployment}} -n {{namespace}} -p '{"spec":{"minReplicas":{{replicas}}}}'
kubectl scale deployment/{{deployment}} -n {{namespace}} --replicas={{replicas}}
When an HPA manages the deployment, raise its minReplicas with the patch command and skip kubectl scale, because the HPA overrides manual scaling. Use kubectl scale only when no HPA exists.
Verify
Rerun the query from step 1 every minute for 10 minutes. The ratio should drop below threshold and stay there. Then confirm the pods and health endpoint:
kubectl get pods -n {{namespace}} -l {{selector}}
for i in $(seq 1 10); do curl -s -o /dev/null -w '%{http_code}\n' "{{service_url}}/healthz"; sleep 3; done
You should see all 200s and no restarts.
Roll back or escalate
- If a rollback made things worse, run
kubectl rollout undo deployment/{{deployment}} -n {{namespace}}with no revision. That returns the deployment to the revision that was running just before your rollback. - If you scaled out, put
minReplicasback to its original value once traffic is normal. - Escalate to the service owner and incident commander if errors stay above threshold 15 minutes after mitigation, if the cause is a shared dependency, or if data integrity may be affected.
Running this runbook in Runspace
In Runspace, each step above runs inline and its output streams directly underneath. The variables can take values captured from an earlier step’s output, such as the revision you find in step 3. Steps 6 and 7 would be approved by a separate reviewer, and the approval is pinned to the exact command and runbook revision. Each run is recorded in a server-side audit log. Terminal capture can turn the shell session from a real incident into a runbook draft for your team to review. Runspace is in private pilots. Request a pilot with your work email, and we’ll reach out to set up a pilot with your team.
FAQ
What should I check first when 5xx errors spike?
First confirm the error ratio from your metrics. Then check pod health and recent deploys. A deploy that lines up with the start of the errors is the most common cause, and rolling back is usually the fastest fix.
Should I roll back or scale out?
Roll back when errors started right after a deploy, or when logs show new application exceptions. Scale out when errors come from load and the dependencies have headroom. Don't scale when the database connection pool is already saturated.
Why does kubectl scale not stick?
If a HorizontalPodAutoscaler manages the deployment, it overrides manual replica counts. Raise the HPA's minReplicas instead.
Can I run this runbook in Runspace?
Yes. In Runspace each step runs inline with streamed output, steps that change production require approval from a separate reviewer, and every run is recorded in a server-side audit log. Runspace is currently in private pilots.