← Runbook templates

TLS certificate expiry runbook

Paste into Runspace or any Markdown runbook. {{name}} marks a variable.

Use this runbook when a TLS certificate is close to expiry, has already expired, or an alert reports a chain or hostname mismatch. It checks the served certificate with openssl, finds which host or cluster serves it, renews it through certbot or cert-manager, reloads the server, and checks the full chain from the client side.

When to use this

  • A monitor or customer reports certificate has expired or unable to get local issuer certificate.
  • An expiry alert fires inside your warning window, for example 14 days.
  • Automated renewal is supposed to run but the served certificate hasn’t changed.

If the endpoint sits behind a CDN or a managed load balancer that terminates TLS, renew the certificate there instead. Step 2 tells you which case you’re in.

Before you start

  • Access: shell with sudo on the web host (certbot path), or kubectl access to the namespace that owns the Ingress (cert-manager path).
  • Tools: OpenSSL 1.1.1 or later, dig, curl. Certbot 2.x on the host. For Kubernetes, kubectl and cmctl (the cert-manager CLI).
  • Know your ACME limits: Let’s Encrypt allows 5 duplicate certificates for the same set of names per 7 days. Don’t force-renew in a loop.

Variables

  • {{domain}}: hostname on the certificate, for example api.example.com
  • {{host}}: IP or hostname to connect to. Use {{domain}} unless you’re testing one backend directly
  • {{port}}: TLS port, usually 443
  • {{warn_days}}: alert threshold in days, for example 14
  • {{cert_name}}: certbot lineage name, as shown by certbot certificates
  • {{namespace}}: Kubernetes namespace of the Ingress and Certificate
  • {{certificate}}: cert-manager Certificate resource name
  • {{secret}}: TLS Secret referenced by the Ingress

Steps

1. Check what the endpoint is serving

Read the served certificate’s dates, issuer and names without changing anything.

echo | openssl s_client -connect {{host}}:{{port}} -servername {{domain}} 2>/dev/null \
  | openssl x509 -noout -subject -issuer -dates -ext subjectAltName

echo | openssl s_client -connect {{host}}:{{port}} -servername {{domain}} 2>/dev/null \
  | openssl x509 -noout -checkend $(( {{warn_days}} * 86400 ))

Expect: notBefore=/notAfter= lines and a SAN list. The second command prints Certificate will not expire (exit 0) or Certificate will expire (exit 1). Decide: If notAfter is well past the window, the alert may be pointing at a different backend. Repeat this step for each IP from step 2.

2. Find where the certificate is served

Work out whether TLS ends at a CDN, a VM, or a Kubernetes Ingress.

dig +short {{domain}}

# On a candidate VM running nginx
sudo nginx -T 2>/dev/null | grep -nE 'server_name|ssl_certificate'

# In Kubernetes
kubectl get ingress -A -o wide | grep {{domain}}
kubectl get certificate -n {{namespace}}

Expect: the resolved IPs, then either an nginx ssl_certificate path (usually under /etc/letsencrypt/live/) or an Ingress with a matching Certificate. Decide: If the issuer from step 1 belongs to a CDN or cloud provider, stop and renew there. A path under /etc/letsencrypt means the certbot path. A Certificate resource means cert-manager.

3. Inspect renewal state

Find out why automatic renewal didn’t run.

# certbot
sudo certbot certificates
systemctl list-timers | grep -E 'certbot|snap.certbot'
sudo tail -n 50 /var/log/letsencrypt/letsencrypt.log

# cert-manager
kubectl describe certificate {{certificate}} -n {{namespace}}
kubectl get certificaterequest,order,challenge -n {{namespace}}

Expect: certbot lists the lineage, its expiry and file paths. cert-manager shows a Ready condition with a reason. Decide: A missing timer, a failed challenge, or a DNS/HTTP-01 error is the root cause. Fix it before renewing, or the next renewal will fail too.

4. Dry-run the renewal (certbot)

Test issuance against the staging environment.

sudo certbot renew --cert-name {{cert_name}} --dry-run

Expect: Congratulations, all simulated renewals succeeded. Deploy hooks are skipped during a dry run. Decide: If the challenge fails, go to “Roll back or escalate”. Don’t attempt a live renewal.

5. Renew the certificate

Requires approval: requests a new production certificate and counts toward ACME rate limits.

# certbot (renews only if inside the renewal window; add --force-renewal only if needed)
sudo certbot renew --cert-name {{cert_name}}

# cert-manager
cmctl renew {{certificate}} -n {{namespace}}
kubectl wait --for=condition=Ready certificate/{{certificate}} -n {{namespace}} --timeout=300s

Expect: certbot reports the renewed certificate path. kubectl wait prints condition met. Decide: If it times out, go back to step 3 and look at the challenge resources.

6. Reload the server

Requires approval: reloads a production web server.

# nginx on a VM (skip if a certbot deploy hook already reloaded it)
sudo nginx -t && sudo systemctl reload nginx

# Kubernetes: ingress controllers watch the Secret; confirm it holds the new cert
kubectl get secret {{secret}} -n {{namespace}} -o jsonpath='{.data.tls\.crt}' \
  | base64 -d | openssl x509 -noout -dates

Expect: syntax is ok and test is successful, then a clean reload. The Secret shows the new notAfter. Decide: If nginx -t fails, don’t reload. The old config keeps serving.

Verify

Check the full chain and hostname from the client side, against every backend IP from step 2.

openssl s_client -connect {{host}}:{{port}} -servername {{domain}} \
  -verify_return_error -verify_hostname {{domain}} </dev/null 2>&1 \
  | grep -E 'depth=|Verify return code'

curl -sv https://{{domain}} -o /dev/null 2>&1 | grep -E 'expire date|SSL certificate verify'

Expect: Verify return code: 0 (ok), a chain running from your leaf to an intermediate, the new expiry date, and SSL certificate verify ok. If you see unable to get local issuer certificate, the server is sending the leaf without its intermediate. Point nginx at fullchain.pem, not cert.pem.

Roll back or escalate

  • Reload broke the site: certbot keeps earlier versions in /etc/letsencrypt/archive/{{cert_name}}/. Point ssl_certificate at the previous fullchainN.pem/privkeyN.pem, run nginx -t, and reload. This is only useful if the old certificate is still valid.
  • ACME challenges keep failing: escalate to whoever owns DNS or the edge firewall. HTTP-01 needs port 80 reachable, and DNS-01 needs API credentials that work.
  • Rate limited: stop retrying. Serve the current certificate until it expires, or escalate for a certificate from a different issuer.
  • Expired and customer-facing: open an incident and page the service owner while you work through this runbook.

Run this runbook in Runspace

In a wiki, this procedure gets run by copy-pasting commands and nobody records what happened. In Runspace, each step runs inline and its output streams underneath it. {{domain}} and {{cert_name}} are variables, and they can be filled from an earlier step’s output. Steps 5 and 6 need approval from a separate reviewer, and that approval is tied to the exact command and runbook revision. Every run goes into a server-side audit log. Runspace is in private pilots. Request a pilot with your work email and we’ll contact you to set one up with your team.

FAQ

How do I check when a TLS certificate expires from the command line?

Run `echo | openssl s_client -connect host:443 -servername domain 2>/dev/null | openssl x509 -noout -dates`. The notAfter line is the expiry. Add `-checkend <seconds>` to get a pass/fail exit code for alerting.

Why does certbot renew say no renewals were attempted?

Certbot only renews certificates inside their renewal window. If the certificate isn't due yet, nothing happens. Use --force-renewal only when you have to, because it counts toward Let's Encrypt's duplicate-certificate limit of 5 per 7 days.

How do I force cert-manager to renew a certificate?

Run `cmctl renew <certificate> -n <namespace>`, then wait for `kubectl wait --for=condition=Ready certificate/<certificate>` to return. If it doesn't become Ready, inspect the Order and Challenge resources.

Can approvals be required before the renew and reload steps run?

In Runspace, yes. Steps that change production can require approval from a separate reviewer, tied to the exact command and runbook revision, and each run is recorded in a server-side audit log. Runspace is currently in private pilots.