Interviewer: "Your Kafka consumers are up. No errors. The producer is sending at the same rate as always. But the lag keeps growing. Why?"
Candidate: "The consumers are too slow. I would add a few more to the group."
This is not answer, this is reflex.
You didn't ask yourself why the lag is building in first place.
To answer it correctly first you need to know a consumer is watched by the broker using 2 different clocks (settings).
The first is the heartbeat. A background thread sends one every 3 seconds. Miss them for 45 seconds and the consumer is treated as dead.
The second is the poll interval. It measures the gap between two calls for records. The default is 5 minutes.
The heartbeat proves your consumer is alive. The poll proves your consumer is working.
Now the arithmetic.
Consumer polls the Kafka broker and gets 500 records by default. Assume consumer's consumption logic in this case is to call an external service that has always answered in 200 ms.
So that one single consumer can process 500 records in under 2 minutes.
This is working fine for a year.
Today that external service is degraded and takes 2 sec to answer.
Same code. Same 500 records. 16 minutes.
And your consumer logic does not ask for the next batch until it has finished this one.
At minute 5 the poll timer expires. Your consumer does not wait to be thrown out. It takes itself out of the group so somebody else can pick up its partitions.
They go to another consumer. Nothing was committed, so this new consumer fetches the same 500 records, hits the same slow service, and resigns at minute 5 too.
Meanwhile the first consumer finishes all 500 and tries to commit. It fails. It gave those partitions away 11 minutes ago.
The work was done twice. The offset never moved. The lag never comes down.
Now go looking for the error. There isn't one.
Nothing failed, so nothing threw. A background thread logged a warning and left the group on purpose. The only honest signal is that failed commit, and it reads like a routine rebalance.
Processing never reports a problem, because processing is working. It is just working on the same 500 records forever.
A classic consumer group has no delivery-attempt limit and no dead letter queue. Nothing here stops on its own.
You can apply these 3 fixes. Pick by what you know about the work.
- Raise the poll interval to cover your real worst case. A dead consumer then takes longer to spot.
- Take fewer records per poll. 500 down to 20, and even a slow day fits inside the window.
- Keep polling on the main thread with the partition paused, and do the work elsewhere.
The default configuration assumes your work finishes in 5 minutes. Your slowest dependency decides whether that is true.
Adding consumers fixes a throughput problem. The producer never gave you one.