# Kafka Consumer Rebalancing: Assignors & KIP-848

**Kafka consumer rebalancing** is the process by which a consumer group redistributes topic partitions among its members whenever membership or subscribed topics change. A partition assignment strategy decides which member owns which partition. The rebalance protocol (classic eager, cooperative, or the broker-driven KIP-848 protocol) decides how much processing stops while ownership moves.

Rebalancing lets a group add consumers and replace failed ones without moving partitions by hand. While it runs, the partitions being moved are not consumed. Frequent rebalances usually come from poll loops slower than `max.poll.interval.ms`, pods restarting without static membership, or session timeouts too short for the network.

## The partition is the unit of parallelism

Inside one group, a partition is owned by exactly one member at a time. A topic with three partitions supports at most three active members. Extra consumer threads or pods remain idle because there is no partition for them to own.

In a lab on Apache Kafka 4.3.1 (one broker, 3 partitions, 4 members), the fourth member joins, gets a member ID, heartbeats normally, and owns no partition:

```text
$ kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
    --describe --group order-fulfillment --members

GROUP             CONSUMER-ID            HOST        CLIENT-ID       #PARTITIONS
order-fulfillment HpZjnCIcSD6oNs2dn35PhA /127.0.0.1  svc-classic-1   1
order-fulfillment dmxipgLtRYGOfNzVpVstdQ /127.0.0.1  svc-new-1       1
order-fulfillment F9oj7CFqS4uqjak6ao349Q /127.0.0.1  svc-new-3       0
order-fulfillment KAUq69TpT86_Dot41sZzBA /127.0.0.1  svc-new-2       1
```

Consumer assignment is separate from producer partitioning. The producer selects a partition for each record; the consumer group selects an owner for each partition. To go beyond the partition count, add partitions, process records in parallel within a partition when ordering allows it, or use [share groups](https://www.conduktor.io/glossary/kafka-share-groups).

## What triggers a Kafka consumer rebalance

A rebalance starts when:

- A member joins the group, for example a new pod or a restarted consumer without static membership.
- A member leaves cleanly by calling `close()` or `unsubscribe()`. A static member (`group.instance.id`) keeps its slot until the session timeout instead.
- A member stops heartbeating for longer than the session timeout (crash, network partition, long GC pause).
- A member goes longer than `max.poll.interval.ms` between two `poll()` calls and leaves the group on its own.
- A subscribed topic gains partitions, or a regex subscription starts matching a newly created topic.

## Kafka rebalance protocols: eager, cooperative and KIP-848

Kafka has three rebalance protocols. **Eager**: every consumer gives up all its partitions, then the group leader reassigns them. **Cooperative**: consumers give up only the partitions that move, over two rounds, and keep processing the rest. **KIP-848**: the broker computes the assignment and updates each consumer through its heartbeat, so only the partitions that move stop.

![Classic protocol with a group-wide synchronisation barrier versus the KIP-848 protocol where the coordinator reconciles each member through heartbeats](https://www.conduktor.io/assets/images/glossary/kafka-consumer-rebalancing-0.webp)

| | Classic eager | Classic cooperative (KIP-429) | Consumer protocol (KIP-848) |
|---|---|---|---|
| Since | Kafka 0.9 | Kafka 2.4 | GA in Kafka 4.0 |
| Who computes | One member elected leader | One member elected leader | Group coordinator (broker) |
| Sync barrier | Yes, every member rejoins | Yes, two rounds | No |
| What stops | All partitions, all members | Only partitions that move | Only partitions that move |
| RPCs | JoinGroup, SyncGroup, Heartbeat, LeaveGroup | Same | ConsumerGroupHeartbeat |
| Assignor runs | In the client (`partition.assignment.strategy`) | In the client | In the broker (`group.remote.assignor`) |

### Classic eager

On any membership change, every member revokes all partitions and rejoins the group. One member computes the new assignment, and no member processes records until every member has rejoined and received its partitions. One slow or failing member pauses the whole group.

### Classic cooperative (KIP-429)

A member elected leader still computes the assignment, in two rounds. In the first, members give up only the partitions that move. In the second, those partitions go to their new owners. Unchanged partitions keep running.

### Consumer protocol (KIP-848)

The group coordinator on the broker computes the assignment, and no member waits for the others. Each member heartbeats to the coordinator, which computes a target assignment and moves partitions one member at a time. A new owner receives a partition only after the old owner has revoked it. Because the broker holds the assignment, `--describe --members --verbose` shows each member's current and target assignment.

Kafka Streams applications do not use `group.protocol=consumer`. Streams has its own broker-side rebalance protocol, defined in KIP-1071, and the classic deprecation plan in KIP-1274 covers `KafkaConsumer` only.

## Partition assignment strategies compared

The assignor decides which member gets which partition. On the classic protocol it is chosen client-side through `partition.assignment.strategy`, and the group uses the first strategy in the list that every member supports. On the consumer protocol the broker owns the list (`group.consumer.assignors`) and a client can only state a preference through `group.remote.assignor`.

| Assignor | Protocol | Balance | Movement on change | Notes |
|---|---|---|---|---|
| `RangeAssignor` | Classic, eager | Per topic: partitions sorted, split into contiguous ranges | Recomputed from scratch | Default. Co-partitions topics with equal partition counts (same partition number to the same member). Skews when many topics have few partitions |
| `RoundRobinAssignor` | Classic, eager | Across all topics, one partition at a time | Recomputed from scratch | Best spread with many small topics; no co-partitioning |
| `StickyAssignor` | Classic, eager | Balanced across all topics | Minimal, but everything is still revoked first | Stickiness only shapes the next assignment |
| `CooperativeStickyAssignor` | Classic, cooperative | Same logic as sticky | Minimal, and unchanged partitions never stop | Every member must support it; one rolling restart from the default list, two from an eager-only list |
| `range` (server) | Consumer | Per-topic ranges, as above | Sticky by design | Keeps co-partitioning for stream joins |
| `uniform` (server) | Consumer | Balanced across all topics | Sticky by design | Broker default |

Classic consumers use range by default, not cooperative sticky. The default list is `RangeAssignor, CooperativeStickyAssignor`, and the group chooses the first strategy supported by every member. Because the cooperative assignor is already in that list, one rolling restart that removes `RangeAssignor` switches a default-configured group to cooperative rebalancing. Members that list only eager assignors need two: first add `CooperativeStickyAssignor`, then remove the eager one.

Range gives up balance to keep the same partition number of every topic on the same consumer. In the Kafka javadoc example, two consumers on two topics of 3 partitions each end up with four partitions (t0p0, t0p1, t1p0, t1p1) and two. Stream joins need that co-location.

Sticky assignors keep each partition on its current owner whenever balance allows, so fewer consumers drop their buffers and caches when a member joins or leaves. Both KIP-848 server-side assignors are sticky by design.

## Kafka rebalance timeouts: session, heartbeat and max.poll.interval.ms

On the classic protocol, three client settings decide when a member is removed from the group:

- `session.timeout.ms` (default 45000): how long the coordinator waits without a heartbeat before declaring the member dead. Heartbeats come from a background thread, so this detects crashes and network partitions, not slow code.
- `heartbeat.interval.ms` (default 3000): how often that thread heartbeats, usually a third of the session timeout.
- `max.poll.interval.ms` (default 300000): the maximum gap between two calls to `poll()`. This one detects slow processing. Exceed it and the consumer leaves the group on its own, triggering a rebalance while it is still busy.

Under `group.protocol=consumer`, session and heartbeat timing move to the broker. The client settings `heartbeat.interval.ms`, `session.timeout.ms`, and `partition.assignment.strategy` are not supported: setting any of them makes the consumer fail at startup with a `ConfigException`, and the `enforceRebalance()` API is not supported. The broker-wide defaults are `group.consumer.session.timeout.ms` (45000) and `group.consumer.heartbeat.interval.ms` (5000, up from the classic client's 3000), and the per-group overrides are `consumer.session.timeout.ms` and `consumer.heartbeat.interval.ms`, set with `kafka-configs.sh`. `max.poll.interval.ms` and static membership through `group.instance.id` still apply.

A minimal consumer configuration for the new protocol:

```properties
bootstrap.servers=kafka-1.internal:9092
group.id=order-fulfillment
group.protocol=consumer
# preference only; the broker's group.consumer.assignors list wins
group.remote.assignor=uniform
# still enforced client-side, sent as the rebalance timeout
max.poll.interval.ms=300000
# survives restarts without a rebalance, on either protocol
group.instance.id=order-fulfillment-pod-2
enable.auto.commit=false
# do not set session.timeout.ms, heartbeat.interval.ms or
# partition.assignment.strategy here: they raise a ConfigException
```

## Migrating a consumer group from classic to KIP-848

Classic and consumer-protocol members can share one group during a rolling upgrade. When the first consumer-protocol member joins, the coordinator converts the group and translates the JoinGroup, SyncGroup and Heartbeat calls of the remaining classic members. It converts back when the last consumer-protocol member leaves. The broker config `group.consumer.migration.policy` governs both directions: the default `bidirectional` allows both, `upgrade` and `downgrade` allow one, and `disabled` allows neither. A lab with two consumer-protocol members and one classic member in the same group on Kafka 4.3.1:

```text
$ kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
    --describe --group order-fulfillment --members --verbose

GROUP             CLIENT-ID      #PARTITIONS  CURRENT-EPOCH  CURRENT-ASSIGNMENT  TARGET-EPOCH  TARGET-ASSIGNMENT  UPGRADED
order-fulfillment svc-classic-1  1            4              orders:1            4             orders:1           false
order-fulfillment svc-new-1      1            4              orders:2            4             orders:2           true
order-fulfillment svc-new-2      1            4              orders:0            4             orders:0           true

$ kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
    --describe --group order-fulfillment --state

GROUP             COORDINATOR (ID)   ASSIGNMENT-STRATEGY  STATE   #MEMBERS
order-fulfillment localhost:9092 (1) uniform              Stable  3
```

The classic member shows `UPGRADED false` and is placed by the broker's `uniform` assignor like everyone else. When the two consumer-protocol members were stopped, the same group reported type `Classic`, strategy `range`, and the surviving member owning all 3 partitions. `kafka-consumer-groups.sh --list --type consumer` (or `--type classic`) lists which groups are on which protocol. Online conversion requires a classic assignor that does not embed custom metadata in its subscription.

The classic protocol is being removed from `KafkaConsumer` on this schedule (KIP-1274):

| Release | Change |
|---|---|
| Kafka 4.3 (May 2026) | `KafkaConsumer` logs a message when started with the classic protocol. Default unchanged |
| Kafka 5.0 | Default becomes `consumer`; classic support deprecated |
| Kafka 6.0 | Classic support removed; `group.protocol=classic` raises a `ConfigException` |

The consumer protocol does not support client-side assignors, so a custom `partition.assignment.strategy` has no direct equivalent. The Kafka 4.3 documentation also lists rack-aware assignment as not yet fully supported.

## Kafka rebalance storms: causes, duplicates and diagnosis

A rebalance storm is a group that cannot stay `Stable` because each rebalance triggers another. Common causes are:

- **Processing slower than `max.poll.interval.ms`.** The member leaves, its partitions move, the new owner hits the same slow records and leaves too.
- **Rolling deployments without static membership.** Every pod restart is a leave and a join. Ten pods deployed one at a time is twenty rebalances.
- **Broker restarts or coordinator moves.** The new coordinator reloads group state from `__consumer_offsets` and members rediscover it. A rebalance follows only for members whose heartbeats do not reach it within the session timeout, so slow failovers and short session timeouts turn a broker restart into a rebalance.
- **Autoscaling on consumer lag.** The rebalance raises lag, and the autoscaler answers with another scale-up.

A rebalance does not remove anything from the log. The new owner resumes from the last committed offset, so records processed but not yet committed run again: the usual symptoms are duplicates and a temporary lag spike. Records are skipped only when an offset was committed before its record was fully processed, for example with auto-commit while processing is handed to another thread.

To limit those duplicates, prefer an explicit `commitSync()` in `onPartitionsRevoked()` over auto-commit. On the classic protocol the consumer does send an auto-commit before it rejoins, but only on a best-effort basis: it waits at most the rebalance timeout, and the rebalance goes ahead if the commit fails ([why auto-commit is not as safe as it looks](https://www.conduktor.io/blog/kafka-production-pitfalls)). The [rebalance listener walkthrough](https://www.conduktor.io/kafka/java-consumer-rebalance-listener) shows the pattern. Neither path helps a member that was already removed for exceeding `max.poll.interval.ms`: its commit fails with `CommitFailedException`, and the new owner reprocesses.

A rebalance adds lag on its own: partitions stop, buffers are dropped, the new owner starts cold. An autoscaler that measures that same lag will answer the spike it caused with another scale-up.

![Feedback loop between a lag-based autoscaler and the group coordinator: lag rises, replicas are added, a rebalance pauses partitions, lag rises again](https://www.conduktor.io/assets/images/glossary/kafka-consumer-rebalancing-1.webp)

Avoid this feedback loop with a cooldown after scaling, separate scale-up and scale-down thresholds, and a signal that rebalancing does not inflate, such as processing time per record. The thresholds in the diagram (scale up at 2x, down at 0.5x) are illustrative examples, not recommended values.

To stabilise a group, apply these in order of effect:

1. Set a `group.instance.id` that is unique and stable per replica, such as a StatefulSet pod name, so restarts within the session timeout do not rebalance. A Kubernetes Deployment pod gets a new name on every restart, so an ID derived from it makes each restart a new static member. Two live instances must never share one ID: on the classic protocol the coordinator fences one of them with `FENCED_INSTANCE_ID`, a fatal error the client does not recover from.
2. Size `max.poll.records` and `max.poll.interval.ms` from the measured worst-case batch, not from defaults.
3. Move to `group.protocol=consumer`, or at least to `CooperativeStickyAssignor` on classic, so a rebalance stops only the partitions that move. Check non-Java clients separately: librdkafka defaults to `range,roundrobin`, both eager, and does not allow eager and cooperative strategies in the same list, so the Java two-step migration does not apply ([librdkafka vs the Java client](https://www.conduktor.io/blog/librdkafka-vs-java-client)).
4. Autoscale on processing time per record rather than lag, with a cooldown after each scale event.
5. Watch group state, not just lag. A group cycling through `Reconciling` or `PreparingRebalance` is the storm; lag only reports it later.

The native check is `kafka-consumer-groups.sh --describe --state` and `--members --verbose`, as shown above. The Java client exposes the same information as JMX metrics:

- `consumer-coordinator-metrics`: `rebalance-rate-per-hour`, `rebalance-total`, `failed-rebalance-total` and `last-rebalance-seconds-ago`. A rising rate on a group with a stable replica count is the storm.
- `consumer-metrics`: `time-between-poll-max` and `last-poll-seconds-ago`. A `time-between-poll-max` approaching `max.poll.interval.ms` predicts the next poll-timeout leave.

On the classic protocol, a member that exceeds `max.poll.interval.ms` logs a warning starting with `consumer poll timeout has expired`, which names the slow poll loop as the cause.

Tools like Conduktor Console show the same group state and membership continuously, alongside per-partition lag. For a vendor-aware walkthrough of lag operations see [Kafka consumer lag](https://www.conduktor.io/kafka-consumer-lag).

**What triggers a Kafka consumer rebalance?**

Any change in group membership or subscribed metadata: a member joins, leaves, misses its session timeout, exceeds max.poll.interval.ms between polls, a subscribed topic gains partitions, or a regex subscription starts matching a new topic. Under the KIP-848 protocol, a change in the group's subscription or metadata bumps the group epoch and the coordinator computes a new target assignment.

**Does Kafka lose messages during a rebalance?**

No data leaves the log. The new owner of a partition resumes from the last committed offset, so records that were processed but not yet committed are processed again, which shows up as duplicates and a temporary lag spike. Records are skipped only if their offsets were committed before processing finished, for example with auto-commit while processing runs on another thread. Idempotent processing or transactional commits absorb the duplicates.

**Which partition assignment strategy is the default?**

On the classic protocol the client default is the list RangeAssignor, CooperativeStickyAssignor, and the group uses range because it is the first strategy every member supports. On the KIP-848 protocol the broker default assignor is uniform, with range available for co-partitioned joins.

**Is cooperative sticky the default rebalancing in Kafka?**

No. CooperativeStickyAssignor is second in the default list so that upgrades work, but a group of default-configured consumers still uses range under the eager protocol. To get cooperative rebalancing on classic, set partition.assignment.strategy to CooperativeStickyAssignor on every member: one rolling restart if members run the default list, two if they list only eager assignors.

**Which consumer settings stop working with group.protocol=consumer?**

session.timeout.ms, heartbeat.interval.ms and partition.assignment.strategy cannot be set: the consumer fails at startup with a ConfigException. The enforceRebalance() API is not supported either. Session timeout and heartbeat interval become broker-side settings (group.consumer.session.timeout.ms and group.consumer.heartbeat.interval.ms, overridable per group as consumer.session.timeout.ms and consumer.heartbeat.interval.ms). max.poll.interval.ms and group.instance.id still apply.

**Can classic and KIP-848 consumers share one group?**

Yes. When the first consumer-protocol member joins, the coordinator converts the group and translates the classic members' JoinGroup, SyncGroup and Heartbeat calls. When the last consumer-protocol member leaves, the group converts back to classic. The classic assignor must not embed custom metadata for the conversion to work.

**When will the classic consumer protocol be removed?**

Under KIP-1274, Kafka 4.3 warns when a consumer starts with classic, Kafka 5.0 switches the default to consumer and deprecates classic, and Kafka 6.0 removes classic support from KafkaConsumer.

## Related Pages

- [Kafka Consumer Groups Explained](https://www.conduktor.io/glossary/kafka-consumer-groups-explained): the group model, offsets and coordinator that rebalancing operates on.
- [Kafka Partitioning Strategies](https://www.conduktor.io/glossary/kafka-partitioning-strategies-and-best-practices): the producer-side choice of partition per record, often confused with consumer assignment.
- [Kafka Partitions Explained](https://www.conduktor.io/glossary/kafka-partitions-explained): why partition count is the parallelism ceiling.
- [Consumer Lag Monitoring](https://www.conduktor.io/glossary/consumer-lag-monitoring): reading the lag spike a rebalance leaves behind.
- [Kafka Share Groups](https://www.conduktor.io/glossary/kafka-share-groups): the consumption model that removes the one-owner-per-partition rule.
- [What's Really Inside Kafka's poll()?](https://www.conduktor.io/blog/kafka-consumer-poll-explained): the heartbeat thread, coordinator connections and the group error codes a consumer has to handle, including `FENCED_INSTANCE_ID`.
- [Incremental Rebalance and Static Group Membership](https://www.conduktor.io/kafka/consumer-incremental-rebalance-and-static-group-membership): configuration walkthrough for cooperative rebalancing and group.instance.id.

## Sources and References

- [Consumer Rebalance Protocol (Apache Kafka 4.3 documentation)](https://kafka.apache.org/43/operations/consumer-rebalance-protocol/)
- [KIP-848: The Next Generation of the Consumer Rebalance Protocol](https://cwiki.apache.org/confluence/display/KAFKA/KIP-848%3A+The+Next+Generation+of+the+Consumer+Rebalance+Protocol)
- [KIP-429: Kafka Consumer Incremental Rebalance Protocol](https://cwiki.apache.org/confluence/display/KAFKA/KIP-429%3A+Kafka+Consumer+Incremental+Rebalance+Protocol)
- [KIP-1274: Deprecate and remove support for Classic rebalance protocol in KafkaConsumer](https://cwiki.apache.org/confluence/display/KAFKA/KIP-1274%3A+Deprecate+and+remove+support+for+Classic+rebalance+protocol+in+KafkaConsumer)
- [Apache Kafka 4.3.0 Release Announcement](https://kafka.apache.org/blog/2026/05/22/apache-kafka-4.3.0-release-announcement/)
- [Consumer Configs (Apache Kafka 4.3 documentation)](https://kafka.apache.org/43/configuration/consumer-configs/)
- [Group Configs (Apache Kafka 4.3 documentation)](https://kafka.apache.org/43/configuration/group-configs/)
- [RangeAssignor (Apache Kafka 4.3 Javadoc)](https://kafka.apache.org/43/javadoc/org/apache/kafka/clients/consumer/RangeAssignor.html)
