How to Manage Kafka ACLs and Security at Scale

Ron Kapoor September 29, 2026 7 min read
An isometric wireframe on a dark teal starfield. On the left, an access request card lists a person, a checkmark, a document and a clock. Those four icons trail off on dashed lines, each struck with an x, while a single arrow carries the request into a Kafka broker cube topped with a padlock shield. A second arrow leads to a card on the right with a person, a book and the Kafka logo.

When the analytics team needs to read orders, they usually open a ticket, and a platform engineer grants the access with kafka-acls --add. The team starts consuming, and the ticket is closed.

That ACL tends to stay long after the reasons for it are gone. The engineer who added it may have moved teams, and the report that needed the data may have been retired. The broker still lets User:analytics-etl read orders, and nobody can say whether it still should.

That's what managing Kafka ACLs at scale usually looks like, and the fix isn't a faster way to write them. Past a few dozen teams, ACLs should be an output: generated from who owns each topic and what that owner approved, not typed in by hand.

An ACL is a precise rule, and it works

An ACL tells the broker that a principal may, or may not, perform one operation on one resource. The broker checks ACLs on every request, and with allow.everyone.if.no.acl.found=false it refuses anything that no allow rule covers.

Each ACL has five parts:

FieldExampleWhat it records
PrincipalUser:analytics-etlWho
OperationReadWhat they can do
ResourceTopic orders, literal or prefixedOn what
Host*From where
PermissionAllow or DenyAllowed or refused
That's a good enforcement primitive. It's small, the broker evaluates it right next to the data, and a prefixed pattern lets one rule cover every topic that starts with orders.. Teams on Strimzi declare ACLs on a KafkaUser resource and let the operator apply them, and plenty of others keep their ACLs in Terraform. With that kind of automation, ACLs hold up fine for a handful of teams.

The ACL keeps the grant and drops the decision

Look at the table again. There's no field for who asked, who approved, why, or until when. The request carried all four, and the ACL kept none of them.

access requestrequesteranalytics teamapproverorders ownerreasonrevenue reportuntil2026-12-31whoanalytics-etlwhatRead on orderskafka-acls --addKafka ACLUser:analytics-etlRead ยท Topic:ordersrequester, approver, reason, until:not kept
The request knew who asked, who agreed, why and for how long. The ACL keeps who and what.

Most of what makes ACLs hard at scale comes from that gap:

  • Nobody can safely remove one. Deleting an ACL means proving that nothing still depends on it, and the ACL gives you nothing to start from. So it stays.
  • "Who can read this topic?" is hard to answer. A topic can be covered by a literal ACL, a prefixed one and a wildcard one at the same time, and the answer is all three combined, on every cluster.
  • The wrong person approves. The platform engineer running the command can check the syntax. Whether analytics should see order data is a question for the team that owns orders, and that team often isn't asked.
  • Every grant runs through a few people. Adding an ACL needs ALTER on the cluster, so the handful of engineers who hold that permission end up processing every request.

In the conversations we have with platform teams, two things come up again and again. Small central teams, sometimes four people, still apply access by hand from tickets. And security reviews keep asking a question the cluster can't answer: when was this team given access to that topic, and who made the call?

"But our ACLs are in Git"

๐Ÿšซ "Our ACLs are in Git, so they're managed."

Keeping ACLs in Git helps a lot, and most teams we talk with won't give up Terraform or GitOps for access, nor should they. You get a history, a diff and a reviewer on every change.

But look at who writes the pull request and who reviews it. Usually a platform engineer writes it for the requester, and another platform engineer approves it. The commit records who merged the line. It doesn't record that the owner of orders agreed, because the owner wasn't part of the review, and nothing in the repository says who that owner is.

So Git fixes the audit trail for the ACL, not the decision behind it. The platform team is still the queue for every grant, and it's still approving requests it has no basis to judge.

Record ownership and grants, then generate the ACLs

The fix is to keep two records that people keep up to date, and treat ACLs as something built from them. The first is ownership: the orders team owns every topic that starts with orders.. The second is grants: the orders team approved analytics reading orders.created for the revenue report.

A generator turns those records into ACLs on the broker. Anything written by hand bypasses both records, which is exactly the problem you're trying to remove:

ownershiporders.* โ†’ orders teamapproved grantsanalytics reads orders.createdgenerateACLson the brokerkafka-acls --add by handno owner, no reason
Generated ACLs trace back to an owner and a grant. Hand-written ones trace back to nobody.

With both records in place, the hard questions get short answers. The list of who can read orders.created is the list of grants on it. Revoking access means deleting a grant, and the generator removes the ACL. The platform team sets the rules, like naming and which grants need a second review, and stops approving individual requests.

People need the same split, with a different enforcement point:

ApplicationsPeople
IdentityOne service account per applicationTheir SSO account and IdP groups
Access comes fromGrants approved by the topic ownerGroup membership mapped to roles
Enforced byACLs on the brokerThe management layer that runs their Kafka calls
Removed whenThe grant or the application is deletedThey leave the group in the IdP

Conduktor generates them from federated ownership

We built this into Conduktor Console as federated ownership. A platform team declares an application instance with its service account and the prefixes it owns. Console then grants that service account READ, WRITE and DESCRIBE_CONFIGS on its topics and READ on its consumer groups. Two application instances can't claim overlapping prefixes on the same cluster, so every topic has one owner.

Access between teams is a request to the owning team, not to the platform team. Once the owner approves, the grant is a resource like this:

apiVersion: self-serve/v1
kind: ApplicationInstancePermission
metadata:
  application: "orders-app"
  appInstance: "orders-app-prod"
  name: "orders-created-to-analytics"
spec:
  resource:
    type: TOPIC
    name: "orders.created"
    patternType: LITERAL
  userPermission: NONE
  serviceAccountPermission: READ
  grantedTo: "analytics-app-prod"

metadata names the owning application, and resource is one topic inside its orders. prefix. serviceAccountPermission: READ gives the analytics service account READ and DESCRIBE_CONFIGS ACLs, while userPermission: NONE gives the analytics team's people nothing in the UI.

Every change goes into Git, and the topic catalog shows every application with access to a topic. Platform teams can also attach a policy to grants, for example refusing any WRITE grant, so owners approve within limits the platform sets.

What it doesn't do. Declining a request in Console doesn't remove access that was already applied: revoking is its own step. And ACLs written by hand for other principals aren't covered by those records, so they need a cleanup of their own.

Where to start

You don't need to move every topic at once. An order that works:

  1. Give every application its own service account. A shared principal makes the rest impossible, because a grant can't be traced back to one application.
  2. Record an owner for every prefix. Start with the topics that hold customer or payment data, and let the rest follow.
  3. Send access requests to the owner. The platform team defines the rules, and the owning team says yes or no.
  4. Generate ACLs from approved grants. Keep kafka-acls for emergencies, and write any emergency grant back into the records afterwards.
  5. Then clean up. List the ACLs that no grant explains, find out who they belong to, and remove what nobody claims.

Kafka ACLs are a good way to enforce access and a poor place to keep the decisions behind it. Record who owns each topic and what each owner approved, and let the ACLs be generated from those records and removed with them. That's what keeps access explainable once there are more teams than a platform team can review by hand.

See how Conduktor manages Kafka ACLs โ†’


Related: Kafka Schema Registry: Open by Default โ†’ ยท The Apache Kafka Security Playbook โ†’ ยท Who Produces and Consumes a Kafka Topic? โ†’