# How to Manage Kafka ACLs and Security at Scale

When the analytics team needs to read `orders`, they usually open a ticket, and a platform engineer grants the access with `kafka-acls --add`. The team starts consuming, and the ticket is closed.

That ACL tends to stay long after the reasons for it are gone. The engineer who added it may have moved teams, and the report that needed the data may have been retired. The broker still lets `User:analytics-etl` read `orders`, and nobody can say whether it still should.

That's what managing Kafka ACLs at scale usually looks like, and the fix isn't a faster way to write them. Past a few dozen teams, ACLs should be an output: generated from who owns each topic and what that owner approved, not typed in by hand.

## An ACL is a precise rule, and it works

An [ACL](https://www.conduktor.io/glossary/kafka-acls-and-authorization-patterns) tells the broker that a principal may, or may not, perform one operation on one resource. The broker checks ACLs on every request, and with `allow.everyone.if.no.acl.found=false` it refuses anything that no allow rule covers.

Each ACL has five parts:

| Field | Example | What it records |
|---|---|---|
| Principal | `User:analytics-etl` | Who |
| Operation | `Read` | What they can do |
| Resource | `Topic` `orders`, literal or prefixed | On what |
| Host | `*` | From where |
| Permission | `Allow` or `Deny` | Allowed or refused |

That's a good enforcement primitive. It's small, the broker evaluates it right next to the data, and a prefixed pattern lets one rule cover every topic that starts with `orders.`. Teams on [Strimzi](https://strimzi.io/) declare ACLs on a `KafkaUser` resource and let the operator apply them, and plenty of others keep their ACLs in Terraform. With that kind of automation, ACLs hold up fine for a handful of teams.

## The ACL keeps the grant and drops the decision

Look at the table again. There's no field for who asked, who approved, why, or until when. The request carried all four, and the ACL kept none of them.

The request knew who asked, who agreed, why and for how long. The ACL keeps who and what.

Most of what makes ACLs hard at scale comes from that gap:

- **Nobody can safely remove one.** Deleting an ACL means proving that nothing still depends on it, and the ACL gives you nothing to start from. So it stays.
- **"Who can read this topic?" is hard to answer.** A topic can be covered by a literal ACL, a prefixed one and a wildcard one at the same time, and the answer is all three combined, on every cluster.
- **The wrong person approves.** The platform engineer running the command can check the syntax. Whether analytics should see order data is a question for the team that owns `orders`, and that team often isn't asked.
- **Every grant runs through a few people.** Adding an ACL needs `ALTER` on the cluster, so the handful of engineers who hold that permission end up processing every request.

In the conversations we have with platform teams, two things come up again and again. Small central teams, sometimes four people, still apply access by hand from tickets. And security reviews keep asking a question the cluster can't answer: when was this team given access to that topic, and who made the call?

## "But our ACLs are in Git"

> 🚫 *"Our ACLs are in Git, so they're managed."*

Keeping ACLs in Git helps a lot, and most teams we talk with won't give up Terraform or GitOps for access, nor should they. You get a history, a diff and a reviewer on every change.

But look at who writes the pull request and who reviews it. Usually a platform engineer writes it for the requester, and another platform engineer approves it. The commit records who merged the line. It doesn't record that the owner of `orders` agreed, because the owner wasn't part of the review, and nothing in the repository says who that owner is.

So Git fixes the audit trail for the ACL, not the decision behind it. The platform team is still the queue for every grant, and it's still approving requests it has no basis to judge.

## Record ownership and grants, then generate the ACLs

The fix is to keep two records that people keep up to date, and treat ACLs as something built from them. The first is ownership: the orders team owns every topic that starts with `orders.`. The second is grants: the orders team approved analytics reading `orders.created` for the revenue report.

A generator turns those records into ACLs on the broker. Anything written by hand bypasses both records, which is exactly the problem you're trying to remove:

Generated ACLs trace back to an owner and a grant. Hand-written ones trace back to nobody.

With both records in place, the hard questions get short answers. The list of who can read `orders.created` is the list of grants on it. Revoking access means deleting a grant, and the generator removes the ACL. The platform team sets the rules, like naming and which grants need a second review, and stops approving individual requests.

People need the same split, with a different enforcement point:

| | Applications | People |
|---|---|---|
| Identity | One service account per application | Their SSO account and IdP groups |
| Access comes from | Grants approved by the topic owner | Group membership mapped to roles |
| Enforced by | ACLs on the broker | The management layer that runs their Kafka calls |
| Removed when | The grant or the application is deleted | They leave the group in the IdP |

## Conduktor generates them from federated ownership

We built this into Conduktor Console as [federated ownership](https://www.conduktor.io/federated-ownership). A platform team declares an application instance with its service account and the prefixes it owns. Console then grants that service account `READ`, `WRITE` and `DESCRIBE_CONFIGS` on its topics and `READ` on its consumer groups. Two application instances can't claim overlapping prefixes on the same cluster, so every topic has one owner.

Access between teams is a request to the owning team, not to the platform team. Once the owner approves, the grant is a resource like this:

```yaml
apiVersion: self-serve/v1
kind: ApplicationInstancePermission
metadata:
  application: "orders-app"
  appInstance: "orders-app-prod"
  name: "orders-created-to-analytics"
spec:
  resource:
    type: TOPIC
    name: "orders.created"
    patternType: LITERAL
  userPermission: NONE
  serviceAccountPermission: READ
  grantedTo: "analytics-app-prod"
```

`metadata` names the owning application, and `resource` is one topic inside its `orders.` prefix. `serviceAccountPermission: READ` gives the analytics service account `READ` and `DESCRIBE_CONFIGS` ACLs, while `userPermission: NONE` gives the analytics team's people nothing in the UI.

Every change goes into Git, and the [topic catalog](https://docs.conduktor.io/guide/use-cases/self-service) shows every application with access to a topic. Platform teams can also attach a policy to grants, for example refusing any `WRITE` grant, so owners approve within limits the platform sets.

> **What it doesn't do.** Declining a request in Console doesn't remove access that was already applied: revoking is its own step. And ACLs written by hand for other principals aren't covered by those records, so they need a cleanup of their own.

## Where to start

You don't need to move every topic at once. An order that works:

1. **Give every application its own service account.** A shared principal makes the rest impossible, because a grant can't be traced back to one application.
2. **Record an owner for every prefix.** Start with the topics that hold customer or payment data, and let the rest follow.
3. **Send access requests to the owner.** The platform team defines the rules, and the owning team says yes or no.
4. **Generate ACLs from approved grants.** Keep `kafka-acls` for emergencies, and write any emergency grant back into the records afterwards.
5. **Then clean up.** List the ACLs that no grant explains, find out who they belong to, and remove what nobody claims.

Kafka ACLs are a good way to enforce access and a poor place to keep the decisions behind it. Record who owns each topic and what each owner approved, and let the ACLs be generated from those records and removed with them. That's what keeps access explainable once there are more teams than a platform team can review by hand.

[See how Conduktor manages Kafka ACLs →](https://www.conduktor.io/kafka-acl)

---

Related: [Kafka Schema Registry: Open by Default →](https://www.conduktor.io/blog/kafka-schema-registry-open-by-default) · [The Apache Kafka Security Playbook →](https://www.conduktor.io/blog/apache-kafka-security-playbook) · [Who Produces and Consumes a Kafka Topic? →](https://www.conduktor.io/blog/who-produces-and-consumes-a-kafka-topic)
