← Runbook templates

Redis memory pressure runbook

Paste into Runspace or any Markdown runbook. {{name}} marks a variable.

Use this runbook when a Redis instance is near or at maxmemory, writes fail with OOM command not allowed, eviction counts spike, or the host is swapping. It walks through reading INFO memory, checking the eviction policy, finding big keys, applying a reversible fix, and confirming that memory has stabilized.

When to use this

  • Alerts on used_memory approaching maxmemory, or on rising evicted_keys.
  • Clients log OOM command not allowed when used memory > 'maxmemory'.
  • RSS on the host grows much faster than used_memory, which points to fragmentation.
  • Latency climbs while memory is high. Eviction and large deletes both block the main thread.

This runbook doesn’t cover cluster resharding or capacity planning. If the dataset has simply outgrown the node, escalate.

Before you start

  • Network access to the instance, plus credentials for a user allowed to run INFO, CONFIG GET, MEMORY and SCAN. Changing config also needs CONFIG SET.
  • redis-cli 6.0 or newer on your machine. --memkeys needs 6.0+.
  • If you can, point the scanning steps at a replica. --bigkeys uses SCAN and doesn’t block, but it does add load.
  • On managed services (ElastiCache, MemoryDB, Azure Cache), CONFIG SET is usually blocked. Change those settings through the provider’s parameter group.
  • Export the password once so it stays out of shell history: export REDISCLI_AUTH='...'.

Variables

  • {{redis_host}}: hostname of the primary that’s under pressure.
  • {{redis_port}}: port, usually 6379.
  • {{replica_host}}: a replica used for scanning. Use {{redis_host}} if there’s no replica.
  • {{big_key}}: a key flagged in Step 4.
  • {{original_policy}}: the maxmemory-policy value captured in Step 2.
  • {{original_maxmemory}}: the maxmemory value in bytes, captured in Step 2.

Steps

1. Read current memory state

Purpose: see how close the instance is to its limit, and whether the gap is data, overhead or fragmentation.

redis-cli -h {{redis_host}} -p {{redis_port}} INFO memory | grep -E '^(used_memory|used_memory_human|used_memory_rss_human|used_memory_peak_human|maxmemory|maxmemory_human|maxmemory_policy|mem_fragmentation_ratio|mem_clients_normal|used_memory_dataset):'

Expected output: one field:value line per field. Compare used_memory with maxmemory.

Decision:

  • mem_fragmentation_ratio above 1.5 while used_memory sits well under the limit: go to Step 6 (fragmentation).
  • mem_clients_normal is large (hundreds of MB): go to Step 5 (client buffers).
  • used_memory is close to maxmemory: continue to Step 2.

2. Check eviction policy and limits

Purpose: record the current settings so you can roll back, and confirm whether Redis is allowed to evict at all.

redis-cli -h {{redis_host}} -p {{redis_port}} CONFIG GET 'maxmemory*'
redis-cli -h {{redis_host}} -p {{redis_port}} INFO keyspace
redis-cli -h {{redis_host}} -p {{redis_port}} INFO stats | grep -E '^(evicted_keys|expired_keys):'

Expected output: maxmemory, maxmemory-policy and the sampling settings, followed by db0:keys=N,expires=M,.... Write down maxmemory as {{original_maxmemory}} and maxmemory-policy as {{original_policy}}.

Decision:

  • noeviction: writes will fail at the limit. Decide whether this data is a cache, which can be evicted, or a store, which can’t.
  • A volatile-* policy with expires far below keys: only keys that have a TTL can be evicted, so the instance behaves almost like noeviction.
  • evicted_keys rising fast under an allkeys-* policy: eviction is working, so look for what’s growing (Step 3).

3. Get a memory breakdown

Purpose: let Redis summarize where memory is going.

redis-cli -h {{redis_host}} -p {{redis_port}} MEMORY DOCTOR
redis-cli -h {{redis_host}} -p {{redis_port}} MEMORY STATS

Expected output: MEMORY DOCTOR gives a plain-language diagnosis, such as high fragmentation or large client buffers. MEMORY STATS breaks memory down into dataset.bytes, clients.normal, replication.backlog and other buckets.

Decision: if dataset.bytes accounts for most of the memory, find the big keys next.

4. Find big keys

Purpose: identify the individual keys that use the most memory.

redis-cli -h {{replica_host}} -p {{redis_port}} --bigkeys -i 0.1
redis-cli -h {{replica_host}} -p {{redis_port}} --memkeys -i 0.1
redis-cli -h {{replica_host}} -p {{redis_port}} MEMORY USAGE {{big_key}} SAMPLES 0
redis-cli -h {{replica_host}} -p {{redis_port}} TTL {{big_key}}

Expected output: --bigkeys lists the biggest key of each type by element count. --memkeys ranks keys by bytes. MEMORY USAGE ... SAMPLES 0 gives the exact byte size of one key. TTL returns -1 when the key has no expiry.

Decision: find the owner of {{big_key}}. An unbounded list, stream or set with no TTL usually means an application bug, and deleting the key only buys time. File a fix with the owning team either way.

5. Check client output buffers

Purpose: find clients, often slow subscribers or MONITOR sessions, that are holding large output buffers.

redis-cli -h {{redis_host}} -p {{redis_port}} CLIENT LIST | awk '{for(i=1;i<=NF;i++) if($i ~ /^omem=/){split($i,a,"="); if(a[2]>10485760) print}}'

Expected output: clients whose omem is above 10 MB. An empty result means buffers aren’t the problem.

Decision: if one client is responsible, killing it is the fastest relief (Step 7).

6. Release memory without deleting data

Purpose: return fragmented or freed memory to the OS.

Requires approval: changes allocator and defrag behavior on the production primary.

redis-cli -h {{redis_host}} -p {{redis_port}} MEMORY PURGE
redis-cli -h {{redis_host}} -p {{redis_port}} CONFIG SET activedefrag yes

Expected output: OK for each command. MEMORY PURGE only works with jemalloc, which is the default on Linux. Active defrag uses CPU in the background.

7. Mitigate pressure

Purpose: apply the smallest change that stops writes from failing. Use one option, not all three.

Requires approval: changes eviction behavior, ends a client connection, or permanently deletes data.

# Option A: allow eviction (only if the data is a cache)
redis-cli -h {{redis_host}} -p {{redis_port}} CONFIG SET maxmemory-policy allkeys-lru

# Option B: kill a client with a runaway output buffer (id from Step 5)
redis-cli -h {{redis_host}} -p {{redis_port}} CLIENT KILL ID {{client_id}}

# Option C: remove a confirmed-disposable big key without blocking
redis-cli -h {{redis_host}} -p {{redis_port}} UNLINK {{big_key}}

Expected output: OK, 1, or the number of keys removed. Prefer UNLINK to DEL for large keys, because DEL frees memory on the main thread and can stall Redis. If you need Option B, add {{client_id}} to your variables.

Verify

redis-cli -h {{redis_host}} -p {{redis_port}} INFO memory | grep -E '^(used_memory_human|maxmemory_human|mem_fragmentation_ratio):'
redis-cli -h {{redis_host}} -p {{redis_port}} INFO stats | grep -E '^evicted_keys:'
redis-cli -h {{redis_host}} -p {{redis_port}} --stat -i 5

Check that used_memory is steady or falling, that evicted_keys has stopped climbing fast, and that clients no longer report OOM errors. Leave --stat running for a few minutes and watch for growth to resume.

Roll back or escalate

Restore the settings you captured in Step 2:

redis-cli -h {{redis_host}} -p {{redis_port}} CONFIG SET maxmemory-policy {{original_policy}}
redis-cli -h {{redis_host}} -p {{redis_port}} CONFIG SET maxmemory {{original_maxmemory}}

CONFIG SET changes don’t survive a restart. Only run CONFIG REWRITE if you mean to keep a change, and also update the config management that owns redis.conf. UNLINK can’t be undone, so restoring a deleted key means restoring from an RDB/AOF backup or having the application rebuild it.

Escalate to the owning team if memory keeps growing after mitigation, if the big keys belong to a system of record, or if the dataset legitimately needs more memory. At that point the fix is a bigger node or sharding.

Running this runbook in Runspace

In Runspace, each step above runs inline and its output streams beneath it. Values such as {{original_policy}} can be captured from the Step 2 output and reused in the rollback. Steps marked “Requires approval” go to a separate reviewer, and the approval is pinned to that exact command and runbook revision. Each run lands in a server-side audit log. Runspace is in private pilots. Request a pilot with your work email and we’ll reach out to set one up with your team.

FAQ

Is redis-cli --bigkeys safe to run in production?

It uses SCAN, so it doesn't block the server, but it does add load. Run it against a replica when you can, and add -i 0.1 to sleep between batches.

What is the difference between --bigkeys and --memkeys?

--bigkeys ranks collections by element count and strings by length. --memkeys ranks keys by actual memory use, measured with MEMORY USAGE. A list with few elements can still be the largest key by bytes.

Should I use DEL or UNLINK to remove a large key?

Use UNLINK. It removes the key right away and frees the memory in a background thread. DEL frees the memory on the main thread and can block Redis while a large key is released.

Why are writes failing even though the eviction policy is volatile-lru?

volatile-* policies only evict keys that have a TTL. If most keys have no expiry, Redis has nothing it can evict and returns OOM errors, the same as noeviction.