Kafka Consumer Groups and Partition Rebalancing Explained
When a Kafka consumer is added, stopped, or falls behind, the group may redistribute its partitions. Understanding Kafka consumer groups and partition rebalancing helps developers and platform engineers distinguish normal membership changes from processing stalls, offset mistakes, and repeated work. This focused guide covers the coordination model and operating choices; for Kafka’s broader storage and event-streaming design, start with the Apache Kafka event streaming guide.
Why Consumer Groups Are Being Talked About
Consumer groups are where Kafka’s horizontal scaling meets application behavior. A service can add consumer instances to process more partitions concurrently, but every membership change may cause reassignment. If that reassignment pauses work, repeats records, or interacts badly with offset commits, the result can be a spike in lag or duplicate business actions.
The details matter during routine deployments as well as incidents. A rolling restart, a slow handler that stops polling, or a temporary network failure can change membership. The Apache Kafka project documents the client and broker behavior, but teams still need to choose an assignment protocol, handle revoked partitions safely, and monitor whether their processing rate is keeping up.
What Is a Kafka Consumer Group?
A consumer group is a set of consumer instances that share the work of reading one or more Kafka topics. For each subscribed topic partition, Kafka assigns ownership to at most one active member of a group at a time. Two different groups can independently read the same records, each with its own progress.
For example, an orders-worker group might divide 12 partitions among three application instances. Each member reads its assigned partitions and tracks its own position. A separate warehouse-loader group can read all 12 partitions for analytics without competing with the order workers.
The group ID is the durable identity for that shared progress. Kafka stores committed offsets for each group and partition, so restarting an application with the same group ID resumes from its saved position. Changing the group ID creates a separate consumer history; it does not move or copy the old group’s offsets.
The Problem Rebalancing Solves
Partition ownership cannot remain fixed when the set of consumers or partitions changes. If a member fails, its partitions need to be assigned to remaining members. When a new member joins, existing work may be redistributed to use the added capacity. A rebalance is the coordination process that updates these ownership assignments.
Without group coordination, applications would need to agree on ownership themselves and recover from stale owners after failures. Kafka’s coordinator provides a shared group state and prevents the group from treating a partition as normally assigned to multiple active members at once. It cannot, however, make application-side work atomic with offset commits. If a handler completes a database write and crashes before committing its Kafka offset, another attempt may repeat that write.
How Group Coordination and Rebalancing Work
Each group has a coordinator broker, selected using the group’s ID and the internal __consumer_offsets topic. Consumers discover that coordinator and use it to join, heartbeat, and report progress. The coordinator maintains membership and group state; assignment logic decides which partitions each member should own.
The classic group protocol
In the classic protocol, members join a generation of the group. The coordinator selects a group leader, which runs the configured partition assignor and proposes an assignment. The coordinator then distributes that assignment to members. Members send heartbeats to show they remain active, and the coordinator advances the group to a new generation when membership or subscription changes require a rebalance.
An eager rebalance revokes the group’s existing assignments before distributing the new ones. This is straightforward, but can temporarily stop consumption across the group, even when only one member changed. An incremental cooperative rebalance revokes only partitions that must move. It can reduce the interruption, although some changes may take more than one rebalance round to settle.
The classic protocol’s commonly used assignment strategies include:
| Strategy | Assignment behavior | Operational trade-off |
|---|---|---|
| Range | Assigns contiguous partition ranges per topic | Simple and predictable, but can be uneven across multiple topics |
| Round robin | Distributes partitions cyclically among members | Often balances counts, but may move more ownership when membership changes |
| Sticky | Tries to balance while preserving existing assignments | Limits unnecessary movement, but still uses eager revocation |
| Cooperative sticky | Preserves assignments and transfers only partitions that need to move | Reduces stop-the-world disruption; requires compatible clients and a careful rollout |
The newer consumer rebalance protocol
KIP-848 describes a newer protocol in which the coordinator manages target assignments and members transition toward them incrementally, rather than relying on a single consumer to calculate a complete assignment and a group-wide synchronization barrier. The proposal’s server-side assignors reached general availability in Apache Kafka 4.0. It is a distinct protocol, not just another classic assignor: confirm support and configuration across the broker and client versions before enabling it. See KIP-848: The Next Generation of the Consumer Rebalance Protocol for its design and migration details.
With a compatible Java client, group.protocol=consumer selects this newer protocol; configure its server-side assignor separately using the broker settings supported by that Kafka release. Do not treat partition.assignment.strategy as interchangeable with this mode.
What triggers a rebalance?
Common triggers include a member joining or leaving, a member failing its liveness checks, a subscription changing, or the topic’s partition metadata changing. In classic mode, a consumer that does not call poll() within max.poll.interval.ms can be considered stuck even if its process is still running. Heartbeat and session timeouts detect a member that is no longer communicating; these are related but distinct signals. The exact timing and supported controls depend on the selected protocol and client version, so use the Kafka consumer configuration reference for the deployed release.
Key Concepts: Assignment, Offsets, and Membership
- Partition assignment: Determines which group member may fetch from each subscribed partition. A group cannot process one partition concurrently through multiple members, so the number of partitions limits active parallelism for that subscription.
- Offset: A position within one partition. A committed offset represents where a group should resume; it does not prove that a business operation completed.
- Offset commit: Recording progress for a group. Committing before processing can lose work after a crash; committing after processing can repeat work if the process crashes before the commit.
- Consumer lag: The difference between the log’s latest offset and the group’s committed or consumed position. Lag can rise because of a rebalance, insufficient processing capacity, slow dependencies, or an unhealthy member.
- Static membership: A stable
group.instance.idcan help avoid reassignment during short, planned restarts. It must be unique per live instance. A failed static member may retain its place until its session expires, delaying reassignment. - Rebalance callbacks: Client callbacks let an application stop work on revoked partitions and prepare newly assigned ones. A callback does not make an external database update and an offset commit transactional.
These distinctions are useful when diagnosing duplicates. Rebalancing changes who should continue reading; the offset commit and the handler’s side effects determine whether a record is repeated or skipped.
Real-World Use Cases
Consumer groups are used in event-driven services, change-data-capture pipelines, stream processing, and batch consumers that need independent replay positions. Multiple groups let separate services process the same topic at different speeds, while adding members to one group can increase parallelism up to the available partition count.
For example, an inventory service may need low-latency processing while an analytics job runs more slowly. They should use distinct group IDs so their progress and failures do not interfere. Within the inventory group, adding instances helps only if there are unassigned partitions and the downstream database can handle the added concurrency.
Practical Guide: Configure and Inspect a Group
For a classic-protocol Java consumer, this minimal properties file disables background commits and opts into cooperative-sticky assignment:
bootstrap.servers=localhost:9092
group.id=orders-worker
enable.auto.commit=false
auto.offset.reset=earliest
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor
auto.offset.reset is used when the group has no valid committed offset; it does not rewind an existing group. With automatic commits disabled, the application must commit only after the records covered by that commit have completed according to its retry and idempotency policy. Do not copy this classic assignment setting into a deployment using the newer consumer protocol without checking that protocol’s configuration.
From a Kafka distribution, inspect offsets and lag with:
bin/kafka-consumer-groups.sh \
--bootstrap-server localhost:9092 \
--describe \
--group orders-worker
The output includes topic partitions, current offsets, log-end offsets, lag, and member information. To inspect group membership and state, use:
bin/kafka-consumer-groups.sh \
--bootstrap-server localhost:9092 \
--describe \
--group orders-worker \
--members
bin/kafka-consumer-groups.sh \
--bootstrap-server localhost:9092 \
--describe \
--group orders-worker \
--state
Compare member counts, partition ownership, lag, and processing latency before changing timeouts or adding replicas. A low lag value at one moment does not rule out repeated short rebalances, and increasing a timeout can make recovery from a truly failed consumer slower. For callback behavior, consult the Kafka ConsumerRebalanceListener API and stop work on revoked partitions before handing them off.
Common Misconceptions
“A rebalance means Kafka lost records.”
A rebalance changes partition ownership; it does not delete the retained log. Records can be processed again if the old owner completed work but did not commit its offset. Preventing harmful duplicates requires idempotent handlers or a transaction strategy that covers the relevant side effect.
“Adding consumers always increases throughput.”
One group member cannot normally own the same partition as another member. Once every partition is assigned, extra members are idle, and excessive membership churn can add coordination overhead. More consumers also increase pressure on databases and downstream services.
“Heartbeats mean the consumer is making progress.”
Heartbeats indicate that the member is communicating with the coordinator; they do not prove that records are being processed quickly. A consumer can heartbeat while making slow progress, or exceed its polling interval while performing long synchronous work. Monitor lag and processing duration alongside group state.
“Cooperative rebalancing eliminates pauses.”
Cooperative assignment can avoid revoking partitions that do not need to move, but moved partitions still require a safe handoff. Application processing, client compatibility, and broker or network delays can still affect throughput during transitions.
Related Articles
- Start with the broader Apache Kafka event streaming guide for topics, partitions, replication, and KRaft.
- See how messaging fits into event-driven microservices.
- Review distributed system failures and network partitions for the failure conditions behind timeouts and recovery.
- Use AsyncAPI for event-driven architectures to document message contracts separately from consumer-group runtime behavior.
Changelog
- Initial publication.

