Runbook for an unreachable EC2 instance: check status checks, security groups, NACLs, console output and SSM, then choose reboot or stop/start and verify.
Runbook template for a full Linux disk: find the cause with df, du, find and lsof +L1, safely vacuum journals and old logs, then verify and prevent.
Runnable runbook for a high 5xx error rate on Kubernetes: confirm impact, check health, deploys, logs, DB pool and upstreams, roll back or scale, verify.
Kafka consumer lag runbook: measure lag with kafka-consumer-groups, find stuck consumers, scale or restart safely, avoid unsafe offset resets, verify.
Runnable Kubernetes rollback runbook: detect a bad rollout, read rollout history, roll back with kubectl rollout undo, verify, and escalate.
PostgreSQL failover runbook template: confirm the primary is down, check replica lag, fence it, promote with pg_promote, repoint clients, verify.
Runbook for Redis memory pressure: read INFO memory, check eviction policy, find big keys with --bigkeys and MEMORY USAGE, mitigate safely, verify.
TLS certificate expiry runbook: check expiry with openssl, find where the cert is served, renew with certbot or cert-manager, reload, and verify the chain.