Kafka Consumer Rebalancing: Assignors & KIP-848
Kafka consumer rebalancing is the process by which a consumer group redistributes topic partitions among its members whenever membership or subscribed topics change. A partition assignment strategy decides which member owns which partition. The rebalance protocol (classic eager, cooperative, or the broker-driven KIP-848 protocol) decides how much processing stops while ownership moves.
Rebalancing lets a group add consumers and replace failed ones without moving partitions by hand. While it runs, the partitions being moved are not consumed. Frequent rebalances usually come from poll loops slower than max.poll.interval.ms, pods restarting without static membership, or session timeouts too short for the network.
The partition is the unit of parallelism
Inside one group, a partition is owned by exactly one member at a time. A topic with three partitions supports at most three active members. Extra consumer threads or pods remain idle because there is no partition for them to own.
In a lab on Apache Kafka 4.3.1 (one broker, 3 partitions, 4 members), the fourth member joins, gets a member ID, heartbeats normally, and owns no partition:
$ kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--describe --group order-fulfillment --members
GROUP CONSUMER-ID HOST CLIENT-ID #PARTITIONS
order-fulfillment HpZjnCIcSD6oNs2dn35PhA /127.0.0.1 svc-classic-1 1
order-fulfillment dmxipgLtRYGOfNzVpVstdQ /127.0.0.1 svc-new-1 1
order-fulfillment F9oj7CFqS4uqjak6ao349Q /127.0.0.1 svc-new-3 0
order-fulfillment KAUq69TpT86_Dot41sZzBA /127.0.0.1 svc-new-2 1 Consumer assignment is separate from producer partitioning. The producer selects a partition for each record; the consumer group selects an owner for each partition. To go beyond the partition count, add partitions, process records in parallel within a partition when ordering allows it, or use share groups.
What triggers a Kafka consumer rebalance
A rebalance starts when:
- A member joins the group, for example a new pod or a restarted consumer without static membership.
- A member leaves cleanly by calling
close()orunsubscribe(). A static member (group.instance.id) keeps its slot until the session timeout instead. - A member stops heartbeating for longer than the session timeout (crash, network partition, long GC pause).
- A member goes longer than
max.poll.interval.msbetween twopoll()calls and leaves the group on its own. - A subscribed topic gains partitions, or a regex subscription starts matching a newly created topic.
Kafka rebalance protocols: eager, cooperative and KIP-848
Kafka has three rebalance protocols. Eager: every consumer gives up all its partitions, then the group leader reassigns them. Cooperative: consumers give up only the partitions that move, over two rounds, and keep processing the rest. KIP-848: the broker computes the assignment and updates each consumer through its heartbeat, so only the partitions that move stop.

| Classic eager | Classic cooperative (KIP-429) | Consumer protocol (KIP-848) | |
|---|---|---|---|
| Since | Kafka 0.9 | Kafka 2.4 | GA in Kafka 4.0 |
| Who computes | One member elected leader | One member elected leader | Group coordinator (broker) |
| Sync barrier | Yes, every member rejoins | Yes, two rounds | No |
| What stops | All partitions, all members | Only partitions that move | Only partitions that move |
| RPCs | JoinGroup, SyncGroup, Heartbeat, LeaveGroup | Same | ConsumerGroupHeartbeat |
| Assignor runs | In the client (partition.assignment.strategy) | In the client | In the broker (group.remote.assignor) |
Classic eager
On any membership change, every member revokes all partitions and rejoins the group. One member computes the new assignment, and no member processes records until every member has rejoined and received its partitions. One slow or failing member pauses the whole group.
Classic cooperative (KIP-429)
A member elected leader still computes the assignment, in two rounds. In the first, members give up only the partitions that move. In the second, those partitions go to their new owners. Unchanged partitions keep running.
Consumer protocol (KIP-848)
The group coordinator on the broker computes the assignment, and no member waits for the others. Each member heartbeats to the coordinator, which computes a target assignment and moves partitions one member at a time. A new owner receives a partition only after the old owner has revoked it. Because the broker holds the assignment, --describe --members --verbose shows each member's current and target assignment.
Kafka Streams applications do not use group.protocol=consumer. Streams has its own broker-side rebalance protocol, defined in KIP-1071, and the classic deprecation plan in KIP-1274 covers KafkaConsumer only.
Partition assignment strategies compared
The assignor decides which member gets which partition. On the classic protocol it is chosen client-side through partition.assignment.strategy, and the group uses the first strategy in the list that every member supports. On the consumer protocol the broker owns the list (group.consumer.assignors) and a client can only state a preference through group.remote.assignor.
| Assignor | Protocol | Balance | Movement on change | Notes |
|---|---|---|---|---|
RangeAssignor | Classic, eager | Per topic: partitions sorted, split into contiguous ranges | Recomputed from scratch | Default. Co-partitions topics with equal partition counts (same partition number to the same member). Skews when many topics have few partitions |
RoundRobinAssignor | Classic, eager | Across all topics, one partition at a time | Recomputed from scratch | Best spread with many small topics; no co-partitioning |
StickyAssignor | Classic, eager | Balanced across all topics | Minimal, but everything is still revoked first | Stickiness only shapes the next assignment |
CooperativeStickyAssignor | Classic, cooperative | Same logic as sticky | Minimal, and unchanged partitions never stop | Every member must support it; one rolling restart from the default list, two from an eager-only list |
range (server) | Consumer | Per-topic ranges, as above | Sticky by design | Keeps co-partitioning for stream joins |
uniform (server) | Consumer | Balanced across all topics | Sticky by design | Broker default |
RangeAssignor, CooperativeStickyAssignor, and the group chooses the first strategy supported by every member. Because the cooperative assignor is already in that list, one rolling restart that removes RangeAssignor switches a default-configured group to cooperative rebalancing. Members that list only eager assignors need two: first add CooperativeStickyAssignor, then remove the eager one. Range gives up balance to keep the same partition number of every topic on the same consumer. In the Kafka javadoc example, two consumers on two topics of 3 partitions each end up with four partitions (t0p0, t0p1, t1p0, t1p1) and two. Stream joins need that co-location.
Sticky assignors keep each partition on its current owner whenever balance allows, so fewer consumers drop their buffers and caches when a member joins or leaves. Both KIP-848 server-side assignors are sticky by design.
Kafka rebalance timeouts: session, heartbeat and max.poll.interval.ms
On the classic protocol, three client settings decide when a member is removed from the group:
session.timeout.ms(default 45000): how long the coordinator waits without a heartbeat before declaring the member dead. Heartbeats come from a background thread, so this detects crashes and network partitions, not slow code.heartbeat.interval.ms(default 3000): how often that thread heartbeats, usually a third of the session timeout.max.poll.interval.ms(default 300000): the maximum gap between two calls topoll(). This one detects slow processing. Exceed it and the consumer leaves the group on its own, triggering a rebalance while it is still busy.
Under group.protocol=consumer, session and heartbeat timing move to the broker. The client settings heartbeat.interval.ms, session.timeout.ms, and partition.assignment.strategy are not supported: setting any of them makes the consumer fail at startup with a ConfigException, and the enforceRebalance() API is not supported. The broker-wide defaults are group.consumer.session.timeout.ms (45000) and group.consumer.heartbeat.interval.ms (5000, up from the classic client's 3000), and the per-group overrides are consumer.session.timeout.ms and consumer.heartbeat.interval.ms, set with kafka-configs.sh. max.poll.interval.ms and static membership through group.instance.id still apply.
A minimal consumer configuration for the new protocol:
bootstrap.servers=kafka-1.internal:9092
group.id=order-fulfillment
group.protocol=consumer
# preference only; the broker's group.consumer.assignors list wins
group.remote.assignor=uniform
# still enforced client-side, sent as the rebalance timeout
max.poll.interval.ms=300000
# survives restarts without a rebalance, on either protocol
group.instance.id=order-fulfillment-pod-2
enable.auto.commit=false
# do not set session.timeout.ms, heartbeat.interval.ms or
# partition.assignment.strategy here: they raise a ConfigException Migrating a consumer group from classic to KIP-848
Classic and consumer-protocol members can share one group during a rolling upgrade. When the first consumer-protocol member joins, the coordinator converts the group and translates the JoinGroup, SyncGroup and Heartbeat calls of the remaining classic members. It converts back when the last consumer-protocol member leaves. The broker config group.consumer.migration.policy governs both directions: the default bidirectional allows both, upgrade and downgrade allow one, and disabled allows neither. A lab with two consumer-protocol members and one classic member in the same group on Kafka 4.3.1:
$ kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--describe --group order-fulfillment --members --verbose
GROUP CLIENT-ID #PARTITIONS CURRENT-EPOCH CURRENT-ASSIGNMENT TARGET-EPOCH TARGET-ASSIGNMENT UPGRADED
order-fulfillment svc-classic-1 1 4 orders:1 4 orders:1 false
order-fulfillment svc-new-1 1 4 orders:2 4 orders:2 true
order-fulfillment svc-new-2 1 4 orders:0 4 orders:0 true
$ kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--describe --group order-fulfillment --state
GROUP COORDINATOR (ID) ASSIGNMENT-STRATEGY STATE #MEMBERS
order-fulfillment localhost:9092 (1) uniform Stable 3 The classic member shows UPGRADED false and is placed by the broker's uniform assignor like everyone else. When the two consumer-protocol members were stopped, the same group reported type Classic, strategy range, and the surviving member owning all 3 partitions. kafka-consumer-groups.sh --list --type consumer (or --type classic) lists which groups are on which protocol. Online conversion requires a classic assignor that does not embed custom metadata in its subscription.
The classic protocol is being removed from KafkaConsumer on this schedule (KIP-1274):
| Release | Change |
|---|---|
| Kafka 4.3 (May 2026) | KafkaConsumer logs a message when started with the classic protocol. Default unchanged |
| Kafka 5.0 | Default becomes consumer; classic support deprecated |
| Kafka 6.0 | Classic support removed; group.protocol=classic raises a ConfigException |
partition.assignment.strategy has no direct equivalent. The Kafka 4.3 documentation also lists rack-aware assignment as not yet fully supported. Kafka rebalance storms: causes, duplicates and diagnosis
A rebalance storm is a group that cannot stay Stable because each rebalance triggers another. Common causes are:
- Processing slower than
max.poll.interval.ms. The member leaves, its partitions move, the new owner hits the same slow records and leaves too. - Rolling deployments without static membership. Every pod restart is a leave and a join. Ten pods deployed one at a time is twenty rebalances.
- Broker restarts or coordinator moves. The new coordinator reloads group state from
__consumer_offsetsand members rediscover it. A rebalance follows only for members whose heartbeats do not reach it within the session timeout, so slow failovers and short session timeouts turn a broker restart into a rebalance. - Autoscaling on consumer lag. The rebalance raises lag, and the autoscaler answers with another scale-up.
A rebalance does not remove anything from the log. The new owner resumes from the last committed offset, so records processed but not yet committed run again: the usual symptoms are duplicates and a temporary lag spike. Records are skipped only when an offset was committed before its record was fully processed, for example with auto-commit while processing is handed to another thread.
To limit those duplicates, prefer an explicit commitSync() in onPartitionsRevoked() over auto-commit. On the classic protocol the consumer does send an auto-commit before it rejoins, but only on a best-effort basis: it waits at most the rebalance timeout, and the rebalance goes ahead if the commit fails (why auto-commit is not as safe as it looks). The rebalance listener walkthrough shows the pattern. Neither path helps a member that was already removed for exceeding max.poll.interval.ms: its commit fails with CommitFailedException, and the new owner reprocesses.
A rebalance adds lag on its own: partitions stop, buffers are dropped, the new owner starts cold. An autoscaler that measures that same lag will answer the spike it caused with another scale-up.

Avoid this feedback loop with a cooldown after scaling, separate scale-up and scale-down thresholds, and a signal that rebalancing does not inflate, such as processing time per record. The thresholds in the diagram (scale up at 2x, down at 0.5x) are illustrative examples, not recommended values.
To stabilise a group, apply these in order of effect:
- Set a
group.instance.idthat is unique and stable per replica, such as a StatefulSet pod name, so restarts within the session timeout do not rebalance. A Kubernetes Deployment pod gets a new name on every restart, so an ID derived from it makes each restart a new static member. Two live instances must never share one ID: on the classic protocol the coordinator fences one of them withFENCED_INSTANCE_ID, a fatal error the client does not recover from. - Size
max.poll.recordsandmax.poll.interval.msfrom the measured worst-case batch, not from defaults. - Move to
group.protocol=consumer, or at least toCooperativeStickyAssignoron classic, so a rebalance stops only the partitions that move. Check non-Java clients separately: librdkafka defaults torange,roundrobin, both eager, and does not allow eager and cooperative strategies in the same list, so the Java two-step migration does not apply (librdkafka vs the Java client). - Autoscale on processing time per record rather than lag, with a cooldown after each scale event.
- Watch group state, not just lag. A group cycling through
ReconcilingorPreparingRebalanceis the storm; lag only reports it later.
The native check is kafka-consumer-groups.sh --describe --state and --members --verbose, as shown above. The Java client exposes the same information as JMX metrics:
consumer-coordinator-metrics:rebalance-rate-per-hour,rebalance-total,failed-rebalance-totalandlast-rebalance-seconds-ago. A rising rate on a group with a stable replica count is the storm.consumer-metrics:time-between-poll-maxandlast-poll-seconds-ago. Atime-between-poll-maxapproachingmax.poll.interval.mspredicts the next poll-timeout leave.
On the classic protocol, a member that exceeds max.poll.interval.ms logs a warning starting with consumer poll timeout has expired, which names the slow poll loop as the cause.
Tools like Conduktor Console show the same group state and membership continuously, alongside per-partition lag. For a vendor-aware walkthrough of lag operations see Kafka consumer lag.
What triggers a Kafka consumer rebalance?
Any change in group membership or subscribed metadata: a member joins, leaves, misses its session timeout, exceeds max.poll.interval.ms between polls, a subscribed topic gains partitions, or a regex subscription starts matching a new topic. Under the KIP-848 protocol, a change in the group's subscription or metadata bumps the group epoch and the coordinator computes a new target assignment.
Does Kafka lose messages during a rebalance?
No data leaves the log. The new owner of a partition resumes from the last committed offset, so records that were processed but not yet committed are processed again, which shows up as duplicates and a temporary lag spike. Records are skipped only if their offsets were committed before processing finished, for example with auto-commit while processing runs on another thread. Idempotent processing or transactional commits absorb the duplicates.
Which partition assignment strategy is the default?
On the classic protocol the client default is the list RangeAssignor, CooperativeStickyAssignor, and the group uses range because it is the first strategy every member supports. On the KIP-848 protocol the broker default assignor is uniform, with range available for co-partitioned joins.
Is cooperative sticky the default rebalancing in Kafka?
No. CooperativeStickyAssignor is second in the default list so that upgrades work, but a group of default-configured consumers still uses range under the eager protocol. To get cooperative rebalancing on classic, set partition.assignment.strategy to CooperativeStickyAssignor on every member: one rolling restart if members run the default list, two if they list only eager assignors.
Which consumer settings stop working with group.protocol=consumer?
session.timeout.ms, heartbeat.interval.ms and partition.assignment.strategy cannot be set: the consumer fails at startup with a ConfigException. The enforceRebalance() API is not supported either. Session timeout and heartbeat interval become broker-side settings (group.consumer.session.timeout.ms and group.consumer.heartbeat.interval.ms, overridable per group as consumer.session.timeout.ms and consumer.heartbeat.interval.ms). max.poll.interval.ms and group.instance.id still apply.
Can classic and KIP-848 consumers share one group?
Yes. When the first consumer-protocol member joins, the coordinator converts the group and translates the classic members' JoinGroup, SyncGroup and Heartbeat calls. When the last consumer-protocol member leaves, the group converts back to classic. The classic assignor must not embed custom metadata for the conversion to work.
When will the classic consumer protocol be removed?
Under KIP-1274, Kafka 4.3 warns when a consumer starts with classic, Kafka 5.0 switches the default to consumer and deprecates classic, and Kafka 6.0 removes classic support from KafkaConsumer.
Related Pages
- Kafka Consumer Groups Explained: the group model, offsets and coordinator that rebalancing operates on.
- Kafka Partitioning Strategies: the producer-side choice of partition per record, often confused with consumer assignment.
- Kafka Partitions Explained: why partition count is the parallelism ceiling.
- Consumer Lag Monitoring: reading the lag spike a rebalance leaves behind.
- Kafka Share Groups: the consumption model that removes the one-owner-per-partition rule.
- What's Really Inside Kafka's poll()?: the heartbeat thread, coordinator connections and the group error codes a consumer has to handle, including
FENCED_INSTANCE_ID. - Incremental Rebalance and Static Group Membership: configuration walkthrough for cooperative rebalancing and group.instance.id.
Sources and References
- Consumer Rebalance Protocol (Apache Kafka 4.3 documentation)
- KIP-848: The Next Generation of the Consumer Rebalance Protocol
- KIP-429: Kafka Consumer Incremental Rebalance Protocol
- KIP-1274: Deprecate and remove support for Classic rebalance protocol in KafkaConsumer
- Apache Kafka 4.3.0 Release Announcement
- Consumer Configs (Apache Kafka 4.3 documentation)
- Group Configs (Apache Kafka 4.3 documentation)
- RangeAssignor (Apache Kafka 4.3 Javadoc)