Kafka consumer lag runbook
Use this runbook when a Kafka consumer group falls behind: lag alerts fire, downstream data goes stale, or a service reports slow processing. It walks you through measuring lag per partition, telling slow consumers from stuck ones, scaling or restarting them safely, and confirming recovery. Offset resets only come up as a last resort, behind approval.
When to use this
- A consumer lag alert has fired for a group, or downstream systems are showing stale data.
- A consumer deployment was changed, restarted, or scaled, and throughput has dropped.
- A group is stuck rebalancing, or some partitions show no assigned consumer.
Don’t use it for broker-side problems such as under-replicated partitions or offline brokers. Fix the cluster first, then come back to this runbook.
Before you start
- Access: network access to the brokers. If the cluster uses SASL or TLS, you’ll need a client properties file. If consumers run on Kubernetes, you’ll need
kubectlaccess to their namespace. - Tools: Kafka CLI tools (
kafka-consumer-groups.sh,kafka-topics.sh) from a Kafka distribution that matches your cluster or is newer. Some packages ship them without the.shsuffix. - Context: the owning team’s normal lag baseline, and whether the consumer is idempotent. You need to know that second one before anyone touches offsets.
Variables
{{bootstrap}}: broker bootstrap address, e.g.kafka-1.internal:9092{{client_config}}: path to a client properties file with security settings{{group}}: the consumer group ID{{topic}}: the topic the group is lagging on{{namespace}}: Kubernetes namespace of the consumer deployment{{deployment}}: consumer deployment name{{replicas}}: target replica count, at most the topic’s partition count{{pod}}: a specific consumer pod identified as stuck{{reset_datetime}}: ISO-8601 timestamp to reset to, e.g.2026-10-07T09:00:00.000
Steps
1. Measure lag per partition
Get the current offset, log-end offset and lag for every partition the group consumes.
kafka-consumer-groups.sh --bootstrap-server {{bootstrap}} \
--command-config {{client_config}} \
--describe --group {{group}}
Expect: one row per partition, with the columns TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG CONSUMER-ID HOST CLIENT-ID.
Decide: If lag is spread evenly across partitions, the group as a whole can’t keep up, so go to step 4. If lag is concentrated on a few partitions, suspect a stuck consumer or a hot key, and go to step 2. If CONSUMER-ID shows -, that partition has no owner, so check step 3.
2. Tell slow from stuck
Sample lag twice, a minute apart, and see whether committed offsets move.
for i in 1 2; do
date -u
kafka-consumer-groups.sh --bootstrap-server {{bootstrap}} \
--command-config {{client_config}} \
--describe --group {{group}} | grep {{topic}}
sleep 60
done
Expect: CURRENT-OFFSET increases on healthy partitions.
Decide: If an offset hasn’t moved while LOG-END-OFFSET grew, that partition’s consumer is stuck. Note its HOST and CLIENT-ID. If offsets move but lag still grows, the consumers are slow, not stuck.
3. Check group state and membership
Confirm the group is stable and see which member owns which partitions.
kafka-consumer-groups.sh --bootstrap-server {{bootstrap}} \
--command-config {{client_config}} \
--describe --group {{group}} --state
kafka-consumer-groups.sh --bootstrap-server {{bootstrap}} \
--command-config {{client_config}} \
--describe --group {{group}} --members --verbose
Expect: the state is Stable (groups on the newer consumer protocol may briefly show Assigning or Reconciling), along with the member count and each member’s assignment.
Decide: If the group cycles through PreparingRebalance repeatedly, consumers are probably exceeding max.poll.interval.ms or crashing. Check the logs in step 4. If the state is Empty, no consumers are running.
4. Inspect the consumers and partition count
Check consumer health and how far you can scale.
kubectl get pods -n {{namespace}} -l app={{deployment}} -o wide
kubectl logs -n {{namespace}} {{pod}} --since=15m | tail -n 100
kafka-topics.sh --bootstrap-server {{bootstrap}} \
--command-config {{client_config}} \
--describe --topic {{topic}} | head -n 1
Expect: pod status and restarts, recent errors (poison messages, timeouts, downstream failures), and PartitionCount for the topic.
Decide: A single stuck pod: go to step 5. Every replica is busy and replicas are below the partition count: go to step 6. A downstream dependency is failing: escalate to its owner, because scaling won’t help.
5. Restart a stuck consumer
Restarting the stuck member makes the group hand its partitions to healthy members.
Requires approval: restarts a production consumer and triggers a group rebalance.
kubectl delete pod -n {{namespace}} {{pod}}
Expect: a replacement pod starts and the group returns to Stable. If the consumer uses static membership (group.instance.id), the new pod reclaims the same partitions without a full rebalance.
Decide: If the new pod sticks on the same offset, a poison message is the likely cause. Escalate to the owning team rather than skipping the message yourself.
6. Scale consumers
Add consumers to increase parallelism, up to the partition count.
Requires approval: changes production capacity and triggers a rebalance.
kubectl scale deployment/{{deployment}} -n {{namespace}} --replicas={{replicas}}
kubectl rollout status deployment/{{deployment}} -n {{namespace}}
Expect: the new pods join the group. Any replicas beyond the partition count sit idle.
Decide: If lag still grows at the partition-count ceiling, the fix is more partitions or a faster consumer. That’s a change for the owning team, not for this incident.
7. Reset offsets (last resort)
Skipping or replaying data is only acceptable when the owning team accepts the data loss or reprocessing. Offsets can only be reset while the group has no active members, so scale consumers to zero first.
kafka-consumer-groups.sh --bootstrap-server {{bootstrap}} \
--command-config {{client_config}} \
--reset-offsets --group {{group}} --topic {{topic}} \
--to-datetime {{reset_datetime}} --dry-run
Review the planned NEW-OFFSET for each partition with the data owner.
Requires approval: permanently skips or replays messages for every partition listed.
kafka-consumer-groups.sh --bootstrap-server {{bootstrap}} \
--command-config {{client_config}} \
--reset-offsets --group {{group}} --topic {{topic}} \
--to-datetime {{reset_datetime}} --execute
Decide: Never use --all-topics or --to-latest during an incident unless the owner has explicitly approved dropping the backlog. Afterwards, scale consumers back up.
Verify
- Re-run step 1 every few minutes. Total lag should trend down, and no partition should show
-as its consumer. - Step 3 reports
Stable, with the expected member count. - The lag alert clears, and the downstream owners confirm their data is fresh.
Roll back or escalate
- Scaling: return to the previous replica count with
kubectl scale ... --replicas=<previous>once lag is back at baseline. - Restart: there’s nothing to roll back. If the problem returns, collect the pod logs and escalate.
- Offset reset: you can’t undo it cleanly. Before running it, record the dry-run output so you can return to the old offsets with
--to-offsetper partition if needed. - Escalate to the consumer’s owning team for poison messages, a partition-count ceiling, or downstream failures, and to the Kafka platform team for broker or coordinator errors.
Running this in Runspace
In Runspace, each step runs inline and its output streams under the step. Values such as {{pod}} can be captured from an earlier step’s output. Steps marked “Requires approval” are approved by a separate reviewer, pinned to the exact command and runbook revision, and every run is recorded in a server-side audit log. Runspace is in private pilots: request a pilot and we’ll reach out to set one up with your team.
FAQ
How do I check Kafka consumer lag from the command line?
Run kafka-consumer-groups.sh --bootstrap-server <broker> --describe --group <group>. The LAG column shows, for each partition, the log-end offset minus the group's committed offset.
Why can't I reset offsets on my consumer group?
kafka-consumer-groups.sh only resets offsets for a group with no active members. Stop or scale the consumers to zero, run the reset with --dry-run, review it, then run it again with --execute.
Will adding more consumers always reduce lag?
Only up to the topic's partition count. Each partition is consumed by at most one member of a group, so replicas beyond the partition count sit idle.
Is Runspace available today?
Runspace is in private pilots with a small number of teams. Requesting a pilot puts you in line, and we'll contact you about timing and fit.